Tag · 4 repos
kv-cache
Repositories carrying the kv-cache tag.
Open-source manuscript of “Understanding AI Infrastructure: Quantitative Analysis and System Design” by Bojie Li. Starting from hardware constraints and model architectures, it derives LLM inference and training system designs quantitatively; includes the full text, PDF, supporting calculation tools, and experiments.
[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"