Skip to main content

Tag · 10 repos

llama-cpp

Repositories carrying the llama-cpp tag.

Exact tag match10 sample repositories
419
Soulmate-Halo/heterogeneous-gpu-pd-lab

Accelerate sparse, high-memory hosts with dense-compute cards. Hosts include AI Max+ 395, M3 512G, DGX Spark, and others.

PythonLocal LLMs
531
Osmantic/ODS

ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

PythonLocal LLMsAI Agents
2212
nazuna-research-labs/margpa-runtime-llm

An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.

PythonRAG & Knowledge BasesLocal LLMs
3770
altic-dev/FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X exclusive model access! 😉 - https://x.com/fluidvoiceapp

SwiftSpeech RecognitionMobile & Desktop Apps
4024
ciddwd/overlay-translator

An open-source Android real-time screen translation tool that requires no ROOT, suitable for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, and text-to-speech (TTS); translations can be displayed directly on screen. Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and…

KotlinMobile & Desktop AppsLocal LLMs
4322
sbsamarski/pi-compact-plus

A context-compaction extension for the pi coding agent: model rotation, tool-call elision, percent-based triggers, editable summary instructions, and a full settings UI.

TypeScriptAI Coding AssistantsLocal LLMs
4790
rulith-dev/rulith-inference

Rulith Inference, formerly Strix Llama. Serving a 125B MoE model fast on one AMD Strix Halo machine, on Windows: a patched llama.cpp, a manager, and a desktop app.

PythonLocal LLMsMobile & Desktop Apps
4893
Speedstu/CUDA-for-AMD-Windows

Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.

PowerShellLocal LLMsLLM Inference & Serving
5002
julianmb/q38rocm

Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.

PythonLocal LLMsLLM Inference & Serving
5263
htobty/ReadyLLM

Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.

PythonLocal LLMsLLM Inference & Serving