Category · 64 repos
LLM Inference & Serving
Engines, kernels and servers that make models run fast and cheap. Ranked by star velocity over the last 24 hours.
Showing 51–64 of 64
Minimal C++ LLM Qwen/Gemma inference engine
SGLang is a high-performance serving framework for large language models and multimodal models.
Pure-C transformer inference engine for Apple Silicon — wired memory independent of model size (mmap streaming, NEON/Metal)
Faster Whisper transcription with CTranslate2
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
Qwen3.8-Flash-Next (NVFP4) on DGX Spark in vLLM: one Spark 43.9 tok/s with our disk-backed n-gram table patch, staged gather and reduced-vocab MTP draft; TP2 SPEED 53.7; TP4 CONTEXT 9M KV pool. Launchers, patch, 40-prompt harness, KV ledger, credits. SGLang TP2 day-0 lane kept.
A programmable Mixture-of-Models router for heterogeneous LLM inference
A high-throughput and memory-efficient inference and serving engine for LLMs
Run Stable Diffusion on Android Devices with Snapdragon NPU acceleration. Also supports CPU/GPU inference.
Suckless no dependencies CUDA/C++ port of NVIDIA's DVLT, feed it images, get a 3D point cloud + camera poses. NO python. One fast 5MB binary.
The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
Dual NVIDIA DGX Spark cluster operator: TP=2 LLM sering, one-key model switching, Dify protocol shim, wachdog, ops panel
fastllm is a high-performance, backend-independent inference library. Supports tensor-parallel dense models and hybrid MOE inference; GPUs with 10G+ can run full DeepSeek. A dual-socket 9004/9005 server with one GPU runs the original full-precision model at 20tps single-concurrency; INT4 reaches 30tps single and 60+ concurrent.