Skip to main content
Back to trending

2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks

Best single-card local long-context solution for an RTX 2080 Ti 22G (sm_75): KVMem mainline with 262K context and measured 36.6 tok/s (IQ3+MTP+vision), lossless streaming KV backup, and complete source reproduction chain plus build, patch, and test tools.

Translated from Chinese · show original

RTX 2080 Ti 22G(sm_75)单卡本地长上下文最优解:KVMem 主线 262K 上下文 / 36.6 tok/s 实测(IQ3+MTP+视觉),KV 流式无损备档;完整源码复现链 + 构建/补丁/测试工具

  • cuda
  • inference
  • kv-cache
  • kvmem
  • llama-cpp
  • llm
  • local-llm
  • long-context
  • mtp
  • rtx-2080-ti
  • speculative-decoding
  • turing
  • windows
View on GitHub
Stars
6
Forks
1
+ today
+1
Created
1mo

Ranking data as of October 5, 2026 (UTC).

Star History

Today, hour by hour

6 stars
01

Overview

Best single-card local long-context solution for an RTX 2080 Ti 22G (sm_75): KVMem mainline with 262K context and measured 36.6 tok/s (IQ3+MTP+vision), lossless streaming KV backup, and complete source reproduction chain plus build, patch, and test tools. It ranks #2195 on GitTiger, gaining +1 star on October 5, 2026 (UTC).

The project is written in Python and has 1 fork. It was created 1mo ago.

Installation
git clone https://github.com/kael-odin/2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks.git
cd 2080ti-22G-kvmem-qwen3.8-27B-256K-36.6toks
# see README for setup