Category · 91 repos
Local LLMs
Run language models on your own machine: private, offline and on consumer hardware. Ranked by star velocity over the last 24 hours.
Showing 51–91 of 91
SLM Forge enables you to build your own small language models on Apple Silicon.
Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
OpenSuno-YuE2: local music generation studio for the YuE2 model on Apple Silicon and Windows/NVIDIA, with automatic model downloads.
This little tool controls external displays (connected via USB-C/DisplayPort Alt Mode) using DDC/CI on Apple Silicon Macs. Useful to embed in various scripts.
A sleek, Apple TV+-style streaming interface for your own Plex or Jellyfin server. Guided first-run setup, one page per library, private libraries, full web player, invite-only accounts, and an AI movie assistant (Gemini, ChatGPT, Claude, Grok, Ollama, LM Studio). Co-created with Claude (Anthropic).
The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
A local, offline AI dubbing pipeline: takes a Japanese video (anime or live-action) + a short voice sample, and produces an English-dubbed version with the original character's voice cloned, background music/SFX preserved, and dialogue timing aligned to the video .
An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.
Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal.
Nigate: An open-source NTFS utility for Mac. It supports all Mac models (Intel and Apple Silicon), providing full read-write access, mounting, and management for NTFS drives.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Bird frame for Raspberry Pi - real-time bird detection by audio, fully local AI, rendered as real, hand-cut 1800s bird illustrations. On an e-ink panel, a TV, or any screen.
ggml speech-to-text inference for 16+ model families
Desktop app to generate 3D models from images or prompt using local AI — runs entirely on your GPU
Elektron Digitakt and Digitone (mk1) emulator: runs their own firmware behind a clickable front panel, with live audio and MIDI. Apps for Windows and Apple silicon Macs; bring your own .syx. Unofficial.
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
🚀 PR Agent: The Original Open-Source PR Reviewer. This project is not the Qodo free tier.
Get models from Hugging Face, convert them, add adapters, and run them in Ollama
Convert device-only arm64 iOS SDKs into Apple Silicon iOS Simulator-compatible outputs.
Learn Ollama by doing; run large-model deployments on the CPU. Read online at: https://datawhalechina.github.io/handy-ollama/
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Run all of Wikipedia offline next to a local LLM — sovereign, private, low-carbon, verifiable. Kiwix + Ollama + RAG.
Private on-device AI chat for Android — runs any GGUF model locally via llama.cpp with ARM-optimised SIMD. Zero network permissions, encrypted settings, biometric lock, tamper detection. + GPU Acceleration
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X exclusive model access! 😉 - https://x.com/fluidvoiceapp
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
Kitsune-Tales-E4B-JP and -EN: LoRA fine-tunes of Gemma 4 E4B that write original fantasy light-novel fiction in Japanese and English. Synthetic data, SFT + DPO, bootstrap-CI evaluation, a validated LLM judge, safety audit, GGUF builds.
An open-source Android real-time screen translation tool that requires no ROOT, suitable for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, and text-to-speech (TTS); translations can be displayed directly on screen. Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and…
mactop - Apple Silicon Monitor Top
Native local-first macOS speech-to-text, live captions, subtitle editing, offline translation, and on-device Gemma 4 transcript enhancement.
A context-compaction extension for the pi coding agent: model rotation, tool-call elision, percent-based triggers, editable summary instructions, and a full settings UI.
a security scanner for custom LLM applications
Dulus Ai — Agentic AI, Making Gemini web cappable of running bash commands in your terminal! [Gui, Web, Cli, Telegram, 2,000 MCP, 100K Skills . LiteLLM (100+ providers), local models via Ollama, /lang in 34 languages, Mesa Redonda, I create the first utility coin that can be used 100% as AI quota or
TUI academia: math → Edge AI. Neovim keys. Rank Edge Mage se merece.
Fast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Free and open-source.
Rulith Inference, formerly Strix Llama. Serving a 125B MoE model fast on one AMD Strix Halo machine, on Windows: a patched llama.cpp, a manager, and a desktop app.
Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.