Category · 49 repos
LLM Inference & Serving
Engines, kernels and servers that make models run fast and cheap. Ranked by star velocity over the last 24 hours.
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.
LLM inference in C/C++
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
A framework for efficient model inference with omni-modality models
A high-throughput and memory-efficient inference and serving engine for LLMs
A local inference engine for Apple silicon, built around the model.
NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
vLLM Omni backend plugin for diffusion and multimodal generation on AWS Trainium
Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images
GGUF Quantization support for native ComfyUI models
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
GLM-5.3-Flash on 4× NVIDIA CMP 170HX with vLLM: up to 437 tok/s single-user, 841 tok/s at 8 users, 262K context, TP4 or PP4, one-command container install
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
Private AI infrastructure for medical and neuroscience research, operated in Canada: private LLMs, serverless models, JupyterLab and federated learning on the AI4OS stack.
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Models
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
A code generator for array-based code on CPUs and GPUs
Metric depth from one RGB frame and a phone LiDAR: LiDAR-prompted MoGe-3 with real-time on-device models (Core ML + Metal)
Zero-leak online detection of LLM decoding corruption and Free Web Search for your AI agents! Catch repetition loops, language drift and garbage mid-stream, before the user sees a bad token. Works with any OpenAI-compatible API.
The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. The fastest local inference engine in the world.
Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Tensors and Dynamic neural networks in Python with strong GPU acceleration
A minimal GPU design in Verilog to learn how GPUs work from the ground up
The terminal that outlives its window. Shells survive quitting the app, agents can drive other agents, and remote machines work the same way. Pure Rust, GPU-rendered.
LLM inference in C/C++
7-node ESP32-S3 cluster running a 0.4B LLM via 1.58-bit (BitNet) ternary quantization over SPI daisy-chain
Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)
Python microservice for vehicle license plate detection using Ultralytics YOLOv11. It can run on CPU or GPU (CUDA). The API is built with FastAPI and accepts images via HTTP POST, returning bounding-box coordinates, classes, and scores.
Efficient Triton Kernels for LLM Training
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
Fast, flexible LLM inference
A list of free LLM inference resources accessible via API.
Pure C# PP-OCRv6 inference library: hand-written kernels, GEMM, and other operators for multiple platforms, with very high performance, low memory requirements, and high accuracy. Includes a managed ONNX interpreter and does not depend on Paddle Inference, ONNX Runtime, or native OpenCV libraries.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Hopsworks - Data-Intensive AI platform with a Feature Store
EXL3 (exllamav3 trellis) inference on Intel Arc Battlemage: bit-exact ESIMD kernels as a vLLM XPU plugin
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes