Tag · 10 repos
inference
Repositories carrying the inference tag.
A framework for efficient model inference with omni-modality models
A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.
A high-throughput and memory-efficient inference and serving engine for LLMs
Zero-leak online detection of LLM decoding corruption and Free Web Search for your AI agents! Catch repetition loops, language drift and garbage mid-stream, before the user sees a bad token. Works with any OpenAI-compatible API.
Runtime type system for IO decoding/encoding
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
Infrastructure for continually self‑improving agents
Port of OpenAI's Whisper model in C/C++
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.