2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks
Best single-card local long-context solution for an RTX 2080 Ti 22G (sm_75): KVMem mainline with 262K context and measured 36.6 tok/s (IQ3+MTP+vision), lossless streaming KV backup, and complete source reproduction chain plus build, patch, and test tools.
Translated from Chinese · show originalhide original
RTX 2080 Ti 22G(sm_75)单卡本地长上下文最优解:KVMem 主线 262K 上下文 / 36.6 tok/s 实测(IQ3+MTP+视觉),KV 流式无损备档;完整源码复现链 + 构建/补丁/测试工具
- cuda
- inference
- kv-cache
- kvmem
- llama-cpp
- llm
- local-llm
- long-context
- mtp
- rtx-2080-ti
- speculative-decoding
- turing
- windows
- Stars
- 6
- Forks
- 1
- + today
- +1
- Created
- 1mo
Ranking data as of October 5, 2026 (UTC).
Star History
Today, hour by hour
Overview
Best single-card local long-context solution for an RTX 2080 Ti 22G (sm_75): KVMem mainline with 262K context and measured 36.6 tok/s (IQ3+MTP+vision), lossless streaming KV backup, and complete source reproduction chain plus build, patch, and test tools. It ranks #2195 on GitTiger, gaining +1 star on October 5, 2026 (UTC).
The project is written in Python and has 1 fork. It was created 1mo ago.
git clone https://github.com/kael-odin/2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks.git
cd 2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks
# see README for setup