Category · 26 repos
Fine-tuning & Training
Tools for training, fine-tuning and aligning models, from LoRA to reinforcement learning. Ranked by star velocity over the last 24 hours.
The platform for continuously improving AI agents.
Trace mining for post-training. Filter trajectories, draft verifiers, export SFT and RL data.
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, a
[NeurIPS 2026] DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
Millisecond decisions, any domain: a 0.8B open "System 1" model that picks between your options with calibrated probabilities. One base, swappable LoRA adapters, on your own hardware.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
Word-accurate .lrc lyric sync. Paste your lyrics, point it at any audio file, get back an enhanced LRC with per-word timestamps. Powered by whisper.cpp + WhisperX (wav2vec2 forced alignment) + Demucs vocal isolation. Prebuilt for NVIDIA CUDA, Vulkan, and CPU. Manual editor for fine-tuning, live prev
Halo is an open-source framework built by White Circle for training large language and multimodal models
Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
PyTorch Code for Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
Your own AI workspace on your own machines: chat, agents, coding, knowledge bases (RAG) and fine-tuning with local models. Open source (AGPL-3.0).
Fizgig — LoRA & Fine-tune Studio for Klein 9B, Krea 2, MiniMax H3 & Qwen Image 2.1: train, fine-tune, profile, repair and extract LoRAs & LoKRs
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
A project implementing various agentic RL based on the Slime post-training framework
Source code for segformer-pytorch, which can be used to train your own models.
Infrastructure for continually self‑improving agents
Efficient Triton Kernels for LLM Training
HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.
Kitsune-Tales-E4B-JP and -EN: LoRA fine-tunes of Gemma 4 E4B that write original fantasy light-novel fiction in Japanese and English. Synthetic data, SFT + DPO, bootstrap-CI evaluation, a validated LLM judge, safety audit, GGUF builds.
Policy Optimization for Generative Models: diffusion and flow policies, online fine-tuning, and inference-time guidance and planning.
Type-safe, distributed orchestration of agents, ML pipelines, and real-time inference on your k8s — in pure Python with async/await, also other languages (rust, go and ts)
Go ahead and axolotl questions