Skip to main content

Topic · 12 repos

Trending in Local LLM

Repositories tagged Local LLM, ranked by star velocity over the last 24 hours.

737
Ryan-gsq/ninfer-16g-5070ti-5080-5090-qwen3.8-27b-gsq-rco

NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.

C++Local LLMsLLM Inference & Serving
930
ARahim3/mlx-dspark

Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

PythonLocal LLMsLLM Inference & Serving
974
Morrowmake/glm53-flash-cmp170hx-recipe

GLM-5.3-Flash on 4× NVIDIA CMP 170HX with vLLM: up to 437 tok/s single-user, 841 tok/s at 8 users, 262K context, TP4 or PP4, one-command container install

PythonLLM Inference & ServingLocal LLMs
1449
nicedreamzapp/claude-code-local

Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / h

PythonLocal LLMsAI Coding Assistants
1757
razzant/ouroboros

Ouroboros — self-creating AI agent. Born Feb 16, 2026.

PythonAI AgentsMulti-Agent Systems
2147
axoviq-ai/synthadoc

Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.

PythonRAG & Knowledge BasesProductivity & Notes
2212
nazuna-research-labs/margpa-runtime-llm

An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.

PythonRAG & Knowledge BasesLocal LLMs
2537
MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

PythonLocal LLMsFine-tuning & Training
3254
HKUDS/nanobot

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

PythonAI AgentsMCP Servers & Tools
3615
jegly/OfflineLLM

Private on-device AI chat for Android — runs any GGUF model locally via llama.cpp with ARM-optimised SIMD. Zero network permissions, encrypted settings, biometric lock, tamper detection. + GPU Acceleration

KotlinLocal LLMsMobile & Desktop Apps
4790
rulith-dev/rulith-inference

Rulith Inference, formerly Strix Llama. Serving a 125B MoE model fast on one AMD Strix Halo machine, on Windows: a patched llama.cpp, a manager, and a desktop app.

PythonLocal LLMsMobile & Desktop Apps
5263
htobty/ReadyLLM

Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.

PythonLocal LLMsLLM Inference & Serving