Skip to main content
Back to trending

HyperQwen

Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.

  • kv-cache
  • llm-inference
  • local-llm
  • quantization
  • qwen
  • qwen3
  • rtx-3090
  • speculative-decoding
  • vllm
  • consumer-gpu
  • cuda
  • inference-optimization
View on GitHub
Stars
1,860
Forks
255
+ today
+2
Created
1mo

Ranking data as of October 5, 2026 (UTC).

Star History

Today, hour by hour

1,860 stars
02

Overview

Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks. It ranks #564 on GitTiger, gaining +2 stars on October 5, 2026 (UTC).

The project is written in Python and has 255 forks. It was created 1mo ago.

Installation
git clone https://github.com/syv-ai/HyperQwen.git
cd HyperQwen
# see README for setup