Conference
Memory-Efficient Speculative Decoding for Large Language Model Inference on Commodity GPUs
| Title: | Memory-Efficient Speculative Decoding for Large Language Model Inference on Commodity GPUs |
|---|---|
| Authors: | Patel, Tejas Pravinbhai |
| Source: | 2026 IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT) Communication Systems and Network Technologies (CSNT), 2026 IEEE 15th International Conference on. :390-396 Apr, 2026 |
| Relation: | 2026 IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT) |
| Database: | IEEE Xplore Digital Library |
Be the first to leave a comment!