Skip to main content

Tag · 8 repos

gguf

Repositories carrying the gguf tag.

Exact tag match8 sample repositories
737
Ryan-gsq/ninfer-16g-5070ti-5080-5090-qwen3.8-27b-gsq-rco

NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.

C++Local LLMsLLM Inference & Serving
2212
nazuna-research-labs/margpa-runtime-llm

An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.

PythonRAG & Knowledge BasesLocal LLMs
2537
MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

PythonLocal LLMsFine-tuning & Training
2660
handy-computer/transcribe.cpp

ggml speech-to-text inference for 16+ model families

C++Speech RecognitionLocal LLMs
3333
datawhalechina/handy-ollama

Learn Ollama by doing; run large-model deployments on the CPU. Read online at: https://datawhalechina.github.io/handy-ollama/

Jupyter NotebookLocal LLMsRAG & Knowledge Bases
4024
ciddwd/overlay-translator

An open-source Android real-time screen translation tool that requires no ROOT, suitable for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, and text-to-speech (TTS); translations can be displayed directly on screen. Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and…

KotlinMobile & Desktop AppsLocal LLMs
5002
julianmb/q38rocm

Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.

PythonLocal LLMsLLM Inference & Serving
5263
htobty/ReadyLLM

Deploy, monitor, and tune LLM through a visual interface — plus image / video generation on top of ComfyUI.

PythonLocal LLMsLLM Inference & Serving