Skip to main content

Category · 64 repos

LLM Inference & Serving

Engines, kernels and servers that make models run fast and cheap. Ranked by star velocity over the last 24 hours.

Showing 51–64 of 64

51
rye-001/Qwenium

Minimal C++ LLM Qwen/Gemma inference engine

C++LLM Inference & Serving
52
sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

PythonLLM Inference & ServingComputer Vision
53
shiaho777/stratum

Pure-C transformer inference engine for Apple Silicon — wired memory independent of model size (mmap streaming, NEON/Metal)

CLLM Inference & ServingLocal LLMs
54
SYSTRAN/faster-whisper

Faster Whisper transcription with CTranslate2

PythonLLM Inference & ServingSpeech Recognition
55
syv-ai/HyperQwen

Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.

PythonLLM Inference & ServingLocal LLMs
56
thu-nics/C2C

[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"

PythonLLM Inference & ServingMulti-Agent Systems
57
tonyd2wild/Qwen3.8-Flash-Next-NVFP4-DGX-Spark

Qwen3.8-Flash-Next (NVFP4) on DGX Spark in vLLM: one Spark 43.9 tok/s with our disk-backed n-gram table patch, staged gather and reduced-vocab MTP draft; TP2 SPEED 53.7; TP4 CONTEXT 9M KV pool. Launchers, patch, 40-prompt harness, KV ledger, credits. SGLang TP2 day-0 lane kept.

PythonLLM Inference & ServingProductivity & Notes
58
vllm-project/semantic-router

A programmable Mixture-of-Models router for heterogeneous LLM inference

GoLLM Inference & ServingLLMOps & Gateways
59
vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

PythonLLM Inference & Serving
60
xororz/local-dream

Run Stable Diffusion on Android Devices with Snapdragon NPU acceleration. Also supports CPU/GPU inference.

KotlinLLM Inference & ServingImage Generation & Editing
61
yassa9/dvlt.cu

Suckless no dependencies CUDA/C++ port of NVIDIA's DVLT, feed it images, get a 3D point cloud + camera poses. NO python. One fast 5MB binary.

CudaLLM Inference & Serving3D, Avatars & Digital Humans
62
youssofal/MTPLX

The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

PythonLLM Inference & ServingLocal LLMs
63
zhuorudan/dual-spark-cluster

Dual NVIDIA DGX Spark cluster operator: TP=2 LLM sering, one-key model switching, Dify protocol shim, wachdog, ops panel

ShellLLM Inference & ServingSelf-Hosted Apps
64
ztxz16/fastllm

fastllm is a high-performance, backend-independent inference library. Supports tensor-parallel dense models and hybrid MOE inference; GPUs with 10G+ can run full DeepSeek. A dual-socket 9004/9005 server with one GPU runs the original full-precision model at 20tps single-concurrency; INT4 reaches 30tps single and 60+ concurrent.

C++LLM Inference & Serving