Category · 27 repos
Evaluation & Benchmarks
Benchmarks, leaderboards and frameworks for measuring how good models and agents are. Ranked by star velocity over the last 24 hours.
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
The platform for continuously improving AI agents.
Red-teaming and evaluation framework for AI agents, built around an OWASP-inspired ASI01–ASI10 taxonomy
Framework for evaluating and improving agents
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, a
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
Vivaria is METR's tool for running evaluations and conducting agent elicitation research.
[NeurIPS 2026] DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
An evaluation harness for medical/health AI agents — reproduce and cover multiple benchmarks under one scoring discipline
Methodology and deterministic evaluation tools for iterative improvement of skills, repositories and agent workflows, with independent review and regression checks.
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.
An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.
Real-event-anchored benchmark for detecting AI-generated videos in real-world crisis settings.
A bounded, self-correcting long-term memory layer for LLM agents: distills continuous, contradictory project interaction logs into a bounded, self-updating memory store. Bounded, self-correcting long-term memory for LLM agents — candidate pool + value dynamics + tension adjudication + causal-replay evals.
Combine the power of nix-eval-jobs with nix-output-monitor to speed-up your evaluation and building process.
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
FactCircuit: an auditable claim-and-evidence verification loop with immutable provenance, staged checks, bounded retrieval, and reproducible evaluation.
The LLM Evaluation Framework
Supercharge Your LLM Application Evaluations 🚀
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.
Evidence-grounded CV/job matching with hybrid LLM routing, explicit requirement satisfaction, abstention, deterministic guardrails, and held-out evaluation.
[Public preview] Hard human-written reasoning puzzles: 749 answered items in 15 topics (707 scored core), with protocol and scorer.
evals.typesafe.ai workflow code published
A desktop RetroAchievements client: your library, what to chase next, missable warnings, leaderboards and rank in one window. Plug in a console that streams its memory and the same window goes live. The first console that does is a real PlayStation 2, through a patched Open PS2 Loader.