Skip to main content

Tag · 10 repos

inference

Repositories carrying the inference tag.

Exact tag match10 sample repositories
483
vllm-project/vllm-omni

A framework for efficient model inference with omni-modality models

PythonLLM Inference & ServingImage Generation & Editing
502
usamahz/cpu-performance-engineering

A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.

CLocal LLMsSystems, Compilers & Build Tools
503
vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

PythonLLM Inference & Serving
1931
doofzoff/SIMURG

Zero-leak online detection of LLM decoding corruption and Free Web Search for your AI agents! Catch repetition loops, language drift and garbage mid-stream, before the user sees a bad token. Works with any OpenAI-compatible API.

PythonLLM Inference & ServingAI Agents
2173
gcanti/io-ts

Runtime type system for IO decoding/encoding

TypeScripttypescriptvalidationinference
2536
zml/zml

Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild

ZigSystems, Compilers & Build Tools
2616
ai-dynamo/grove

Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling

GoCloud & KubernetesAI Agents
3500
Human-Agent-Society/reef

Infrastructure for continually self‑improving agents

PythonAI AgentsFine-tuning & Training
4259
ggml-org/whisper.cpp

Port of OpenAI's Whisper model in C/C++

C++Speech Recognition
5010
kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++LLM Inference & ServingMachine Learning & Data Science