Tag · 10 repos
llama-cpp
Repositories carrying the llama-cpp tag.
Accelerate sparse, high-memory hosts with dense-compute cards. Hosts include AI Max+ 395, M3 512G, DGX Spark, and others.
ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X exclusive model access! 😉 - https://x.com/fluidvoiceapp
An open-source Android real-time screen translation tool that requires no ROOT, suitable for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, and text-to-speech (TTS); translations can be displayed directly on screen. Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and…
A context-compaction extension for the pi coding agent: model rotation, tool-call elision, percent-based triggers, editable summary instructions, and a full settings UI.
Rulith Inference, formerly Strix Llama. Serving a 125B MoE model fast on one AMD Strix Halo machine, on Windows: a patched llama.cpp, a manager, and a desktop app.
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.