Category · 91 repos
Local LLMs
Run language models on your own machine: private, offline and on consumer hardware. Ranked by star velocity over the last 24 hours.
Showing 1–50 of 91
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.
Local image or text → game-ready 3D on your own machine. Retopology, repaint, rigging and animation helpers, in a web viewer and a CLI. Runs Pixal3D, TRELLIS.2, Hunyuan3D and SF3D. Apple Silicon and NVIDIA. Every result carries its licence.
The sovereign enclosure. A sovereign, open-source, post-quantum alternative to the Palantir platform family; one core platform and four composable products, built and verified against a complete specification.
Wispr Flow for your lips: hold a key, silently mouth words, and they're typed at your cursor. Runs locally on your Mac.
List of Permanent Free LLM API (API Keys)
The open-source AI voice studio. Clone, dictate, create.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
Accelerate sparse, high-memory hosts with dense-compute cards. Hosts include AI Max+ 395, M3 512G, DGX Spark, and others.
A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for Apple silicon.
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.
ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
Open-source text humanization pipeline with every intermediate step published. Two LLM rewrites at temp 1.3, then two hops across different NMT engines. Four documented methodologies you can read, modify, and run locally.
Project NOMAD is an offline-first knowledge and education server. Wikipedia, thousands of books, courses, maps, and optional local AI, all running on hardware you own with no internet required.
Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
An all-in-one, 100% local AI video, image, and music studio. Director mode plans full music videos and short films from a single prompt. Built on the WanGP pipeline. Install via Pinokio.
Bridge local AI coding agents (Claude Code, Cursor, Gemini CLI, Codex) to messaging platforms (Feishu/Lark, DingTalk, Slack, Telegram, Discord, LINE, WeChat Work). Chat with your AI dev assistant from anywhere — no public IP required for most platforms.
A local inference engine for Apple silicon, built around the model.
Kernel-level antidetect browser with the Playwright API you already write. Python + Node SDKs, MCP-ready, unlimited local profiles. Linux x64 + arm64, macOS Intel + Apple Silicon, Windows x64.
Local AI Developer Platform with Integrated Code Editor, Mixture-of-Agents Collaborative Ensemble.
Local AI, native to your Mac. Chat, serve, monitor, and connect MLX models from one macOS app.
NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images
GGUF Quantization support for native ComfyUI models
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
Android-first on-device AI agent harness — exploring everything possible on Android
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
Brazilian version of Project N.O.M.A.D., focused on offline content in Portuguese, maps of Brazil, local AI, and education.
An unofficial userscript for the ChatGPT web app that alerts you when a response is ready or user action is required. It runs locally in your browser with @grant none and is designed for privacy, security, and responsible use: no external requests, conversation-data collection, automated prompt subm
GLM-5.3-Flash on 4× NVIDIA CMP 170HX with vLLM: up to 437 tok/s single-user, 841 tok/s at 8 users, 262K context, TP4 or PP4, one-command container install
High-performance decentralized coordination protocol for multi-agent LLM systems. Plant-inspired signal diffusion with sub-microsecond processing, QUIC P2P networking, and unified LLM backends (Ollama + Claude).
Invoice OCR POC
A free, offline, private AI text-to-speech desktop app built on Rust 🦜
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
the terminal client for LLMs
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / h
Modified llama.cpp to work with Strix Halo + Radeon R9700 as an accelerator, Dual Strix Halo Machines, or both
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
🎬 Integrates seedance2: an open-source local AI short-drama and comic-drama generation tool — complete the process from story to finished video in one place. Data stays on your machine; a highly flexible short-drama workflow platform for AI live-action and AI comic dramas, all handled locally. Open-source local AI short drama maker: story → storyboard → video, fully offline, your data stays…
Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Models
Model export recipes, Python primitives, and Swift runtime utilities for on-device AI
Flowboard — open-source infinite canvas for AI product videos. Drag nodes for model, product, scene, video — connect them, click Generate. Auto-prompts write themselves. Runs locally; powered by Google Flow + Claude CLI.
Your own AI workspace on your own machines: chat, agents, coding, knowledge bases (RAG) and fine-tuning with local models. Open source (AGPL-3.0).
Build a Raspberry Pi wildlife camera step by step — from Python image capture and motion detection to edge AI.
Metric depth from one RGB frame and a phone LiDAR: LiDAR-prompted MoGe-3 with real-time on-device models (Core ML + Metal)
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads.