Skip to main content

Tag · 4 repos

kv-cache

Repositories carrying the kv-cache tag.

Exact tag match4 sample repositories
1695
bojieli/ai-infra-book

Open-source manuscript of “Understanding AI Infrastructure: Quantitative Analysis and System Design” by Bojie Li. Starting from hardware constraints and model architectures, it derives LLM inference and training system designs quantitatively; includes the full text, PDF, supporting calculation tools, and experiments.

PythonLLM Inference & ServingAwesome Lists & Learning
3184
jjiantong/Awesome-KV-Cache-Optimization

[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

PythonLLM Inference & ServingMachine Learning & Data Science
5420
syv-ai/HyperQwen

Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.

PythonLLM Inference & ServingLocal LLMs
5549
thu-nics/C2C

[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"

PythonMulti-Agent SystemsLLM Inference & Serving