Skip to main content

Tag · 8 repos

vllm

Repositories carrying the vllm tag.

Exact tag match8 sample repositories
415
Orchestra-Research/AI-Research-SKILLs

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

TeXAI Coding AssistantsPrompts & Skills
765
aws-neuron/vllm-omni-neuron

vLLM Omni backend plugin for diffusion and multimodal generation on AWS Trainium

PythonLLM Inference & ServingImage Generation & Editing
974
Morrowmake/glm53-flash-cmp170hx-recipe

GLM-5.3-Flash on 4× NVIDIA CMP 170HX with vLLM: up to 437 tok/s single-user, 841 tok/s at 8 users, 262K context, TP4 or PP4, one-command container install

PythonLLM Inference & ServingLocal LLMs
1228
AlirezaAbedinii/CAIOS

Private AI infrastructure for medical and neuroscience research, operated in Canada: private LLMs, serverless models, JupyterLab and federated learning on the AI4OS stack.

PythonLLM Inference & ServingMachine Learning & Data Science
1617
XHToken/Spark-X2.5

Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Models

Local LLMsAI Agents
1931
doofzoff/SIMURG

Zero-leak online detection of LLM decoding corruption and Free Web Search for your AI agents! Catch repetition loops, language drift and garbage mid-stream, before the user sees a bad token. Works with any OpenAI-compatible API.

PythonLLM Inference & ServingAI Agents
3409
nokia-applied-research/AnyJev

Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)

PythonLLM Inference & Serving
5010
kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++LLM Inference & ServingMachine Learning & Data Science