HyperQwen
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
- kv-cache
- llm-inference
- local-llm
- quantization
- qwen
- qwen3
- rtx-3090
- speculative-decoding
- vllm
- consumer-gpu
- cuda
- inference-optimization
- Stars
- 1,860
- Forks
- 255
- + today
- +2
- Created
- 1mo
Ranking data as of October 5, 2026 (UTC).
Star History
Today, hour by hour
02
Overview
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks. It ranks #564 on GitTiger, gaining +2 stars on October 5, 2026 (UTC).
The project is written in Python and has 255 forks. It was created 1mo ago.
Installation
git clone https://github.com/syv-ai/HyperQwen.git
cd HyperQwen
# see README for setup