Skip to main content

Category · 49 repos

LLM Inference & Serving

Engines, kernels and servers that make models run fast and cheap. Ranked by star velocity over the last 24 hours.

1
Niko1221/Strata

Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

C++LLM Inference & ServingAPIs & SDKs
2
magnitudedev/magnitude

Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.

RustLLM Inference & ServingLocal LLMs
3
ggml-org/llama.cpp

LLM inference in C/C++

C++LLM Inference & Serving
4
Orchestra-Research/AI-Research-SKILLs

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

TeXLLM Inference & ServingAI Coding Assistants
5
vllm-project/vllm-omni

A framework for efficient model inference with omni-modality models

PythonLLM Inference & ServingImage Generation & Editing
6
vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

PythonLLM Inference & Serving
7
incoai/splash

A local inference engine for Apple silicon, built around the model.

C++LLM Inference & ServingLocal LLMs
8
Ryan-gsq/ninfer-16g-5070ti-5080-5090-qwen3.8-27b-gsq-rco

NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.

C++LLM Inference & ServingLocal LLMs
9
aws-neuron/vllm-omni-neuron

vLLM Omni backend plugin for diffusion and multimodal generation on AWS Trainium

PythonLLM Inference & ServingImage Generation & Editing
10
fermionresearch/phonon

Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images

PythonLLM Inference & ServingSpeech Recognition
11
leejet/ComfyUI-GGUF

GGUF Quantization support for native ComfyUI models

PythonLLM Inference & ServingLocal LLMs
12
nobodywho-ooo/nobodywho

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

RustLLM Inference & ServingMobile & Desktop Apps
13
ARahim3/mlx-dspark

Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

PythonLLM Inference & ServingLocal LLMs
14
Morrowmake/glm53-flash-cmp170hx-recipe

GLM-5.3-Flash on 4× NVIDIA CMP 170HX with vLLM: up to 437 tok/s single-user, 841 tok/s at 8 users, 262K context, TP4 or PP4, one-command container install

PythonLLM Inference & ServingLocal LLMs
15
NVIDIA/cuda-samples

Samples for CUDA Developers which demonstrates features in CUDA Toolkit

C++LLM Inference & Serving
16
AlirezaAbedinii/CAIOS

Private AI infrastructure for medical and neuroscience research, operated in Canada: private LLMs, serverless models, JupyterLab and federated learning on the AI4OS stack.

PythonLLM Inference & ServingMachine Learning & Data Science
17
Dreamer-Toby/STEPQuant

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

PythonLLM Inference & Serving
18
mlrun/mlrun

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.

PythonLLM Inference & ServingMachine Learning & Data Science
19
XHToken/Spark-X2.5

Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Models

LLM Inference & ServingLocal LLMs
20
bendlang/bend

Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh

TypeScriptLLM Inference & ServingSystems, Compilers & Build Tools
21
inducer/loopy

A code generator for array-based code on CPUs and GPUs

PythonLLM Inference & ServingAI Coding Assistants
22
sergmister/PromptMoGe

Metric depth from one RGB frame and a phone LiDAR: LiDAR-prompted MoGe-3 with real-time on-device models (Core ML + Metal)

PythonLLM Inference & ServingComputer Vision
23
doofzoff/SIMURG

Zero-leak online detection of LLM decoding corruption and Free Web Search for your AI agents! Catch repetition loops, language drift and garbage mid-stream, before the user sees a bad token. Works with any OpenAI-compatible API.

PythonLLM Inference & ServingAI Agents
24
youssofal/MTPLX

The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

PythonLLM Inference & ServingLocal LLMs
25
jundot/omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

PythonLLM Inference & ServingLocal LLMs
26
Comfy-Org/ComfyUI

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. The fastest local inference engine in the world.

PythonLLM Inference & ServingImage Generation & Editing
27
ai-dynamo/grove

Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling

GoLLM Inference & ServingCloud & Kubernetes
28
hiyouga/LlamaFactory

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

PythonLLM Inference & ServingFine-tuning & Training
29
pytorch/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

PythonLLM Inference & ServingMachine Learning & Data Science
30
adam-maj/tiny-gpu

A minimal GPU design in Verilog to learn how GPUs work from the ground up

SystemVerilogLLM Inference & Serving
31
l0ng-ai/tty7

The terminal that outlives its window. Shells survive quitting the app, agents can drive other agents, and remote machines work the same way. Pure Rust, GPU-rendered.

RustLLM Inference & ServingCLI & Developer Tools
32
unslothai/llama.cpp

LLM inference in C/C++

C++LLM Inference & Serving
33
Low-Zi-Hong/ESP32s3-LLM-Cluster

7-node ESP32-S3 cluster running a 0.4B LLM via 1.58-bit (BitNet) ternary quantization over SPI daisy-chain

C++LLM Inference & Serving
34
nokia-applied-research/AnyJev

Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)

PythonLLM Inference & Serving
35
jefu2011/ReconhecimentoPlaca

Python microservice for vehicle license plate detection using Ultralytics YOLOv11. It can run on CPU or GPU (CUDA). The API is built with FastAPI and accepts images via HTTP POST, returning bounding-box coordinates, classes, and scores.

PythonLLM Inference & Serving
36
linkedin/Liger-Kernel

Efficient Triton Kernels for LLM Training

PythonLLM Inference & ServingFine-tuning & Training
37
noonghunna/club-3090

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.

PythonLLM Inference & ServingLocal LLMs
38
EricLBuehler/mistral.rs

Fast, flexible LLM inference

RustLLM Inference & Serving
39
jtig37/free-llm-api-resources

A list of free LLM inference resources accessible via API.

PythonLLM Inference & Serving
40
sdcb/SimdPaddleOCR

Pure C# PP-OCRv6 inference library: hand-written kernels, GEMM, and other operators for multiple platforms, with very high performance, low memory requirements, and high accuracy. Includes a managed ONNX interpreter and does not depend on Paddle Inference, ONNX Runtime, or native OpenCV libraries.

C#LLM Inference & ServingSystems, Compilers & Build Tools
41
FareedKhan-dev/kimi-k3-in-c

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

CLLM Inference & ServingMachine Learning & Data Science
42
Speedstu/CUDA-for-AMD-Windows

Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.

PowerShellLLM Inference & ServingLocal LLMs
43
julianmb/q38rocm

Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.

PythonLLM Inference & ServingLocal LLMs
44
kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++LLM Inference & ServingMachine Learning & Data Science
45
logicalclocks/hopsworks

Hopsworks - Data-Intensive AI platform with a Feature Store

JavaLLM Inference & ServingMachine Learning & Data Science
46
0xSero/exl3xpu

EXL3 (exllamav3 trellis) inference on Intel Arc Battlemage: bit-exact ESIMD kernels as a vLLM XPU plugin

PythonLLM Inference & Serving
47
bentoml/BentoML

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

PythonLLM Inference & ServingMachine Learning & Data Science
48
htobty/ReadyLLM

Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.

PythonLLM Inference & ServingLocal LLMs
49
kserve/kserve

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

GoLLM Inference & ServingCloud & Kubernetes