Memory-Efficient Speculative Decoding for Large Language Model Inference on Commodity GPUs

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Memory-Efficient Speculative Decoding for Large Language Model Inference on Commodity GPUs
Συγγραφείς: Patel, Tejas Pravinbhai
Πηγή: 2026 IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT) Communication Systems and Network Technologies (CSNT), 2026 IEEE 15th International Conference on. :390-396 Apr, 2026
Relation: 2026 IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT)
Βάση Δεδομένων: IEEE Xplore Digital Library
Περιγραφή
ISBN:9798331551780
9798331551797
ISSN:24735655
23297182
DOI:10.1109/CSNT69054.2026.11502509