Skip to main content

Category · 27 repos

Evaluation & Benchmarks

Benchmarks, leaderboards and frameworks for measuring how good models and agents are. Ranked by star velocity over the last 24 hours.

1
ifixai-ai/iFixAi

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

PythonEvaluation & BenchmarksAI Agents
2
overmind-core/overmind

The platform for continuously improving AI agents.

PythonEvaluation & BenchmarksAI Agents
3
AgentSafeLabs/safelabs-eval

Red-teaming and evaluation framework for AI agents, built around an OWASP-inspired ASI01–ASI10 taxonomy

PythonEvaluation & BenchmarksAI Safety & Red Teaming
4
harbor-framework/harbor

Framework for evaluating and improving agents

PythonEvaluation & Benchmarks
5
calmrocks/ai-engineer-notebooks

Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, a

Jupyter NotebookEvaluation & BenchmarksFine-tuning & Training
6
trycua/cua

Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.

RustEvaluation & BenchmarksAI Agents
7
METR/vivaria

Vivaria is METR's tool for running evaluations and conducting agent elicitation research.

TypeScriptEvaluation & Benchmarks
8
OpenDCAI/DataFlex

[NeurIPS 2026] DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models

PythonEvaluation & BenchmarksFine-tuning & Training
9
embeddings-benchmark/mteb

MTEB: State-of-the-art evaluation of embeddings across languages and modalities

PythonEvaluation & BenchmarksRAG & Knowledge Bases
10
thetahealth/mirobody-eval

An evaluation harness for medical/health AI agents — reproduce and cover multiple benchmarks under one scoring discipline

PythonEvaluation & BenchmarksAI Agents
11
DaizeDong/self-evolve

Methodology and deterministic evaluation tools for iterative improvement of skills, repositories and agent workflows, with independent review and regression checks.

PythonEvaluation & BenchmarksAI Agents
12
FailproofAI/failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.

TypeScriptEvaluation & BenchmarksAI Agents
13
nazuna-research-labs/margpa-runtime-llm

An experimental LLM platform for local and cloud inference, chat, RAG, agents and tools, guardrails, security, monitoring, evaluation, LLM-as-a-Judge, and research into original features.

PythonEvaluation & BenchmarksRAG & Knowledge Bases
14
Shuo-Liang-0111/RA-Bench

Real-event-anchored benchmark for detecting AI-generated videos in real-world crisis settings.

PythonEvaluation & BenchmarksAI Safety & Red Teaming
15
1173591564/Dynamics-memory

A bounded, self-correcting long-term memory layer for LLM agents: distills continuous, contradictory project interaction logs into a bounded, self-updating memory store. Bounded, self-correcting long-term memory for LLM agents — candidate pool + value dynamics + tension adjudication + causal-replay evals.

TypeScriptEvaluation & BenchmarksRAG & Knowledge Bases
16
Mic92/nix-fast-build

Combine the power of nix-eval-jobs with nix-output-monitor to speed-up your evaluation and building process.

PythonEvaluation & Benchmarks
17
langfuse/langfuse

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

TypeScriptEvaluation & BenchmarksMonitoring & Observability
18
superwesleyhys-ux/factcircuit

FactCircuit: an auditable claim-and-evidence verification loop with immutable provenance, staged checks, bounded retrieval, and reproducible evaluation.

PythonEvaluation & BenchmarksAI Safety & Red Teaming
19
confident-ai/deepeval

The LLM Evaluation Framework

PythonEvaluation & Benchmarks
20
vibrantlabsai/ragas

Supercharge Your LLM Application Evaluations 🚀

PythonEvaluation & BenchmarksLLMOps & Gateways
21
EverMind-AI/SkillCorpus

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

PythonEvaluation & BenchmarksRAG & Knowledge Bases
22
comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

PythonEvaluation & BenchmarksLLMOps & Gateways
23
Simreal-AI/Simreal-MLBench

[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.

PythonEvaluation & BenchmarksAI Agents
24
jaysanghavi55/fitlens

Evidence-grounded CV/job matching with hybrid LLM routing, explicit requirement satisfaction, abstention, deterministic guardrails, and held-out evaluation.

TypeScriptEvaluation & BenchmarksMachine Learning & Data Science
25
Simreal-AI/hard-puzzle-benchmark

[Public preview] Hard human-written reasoning puzzles: 749 answered items in 15 topics (707 scored core), with protocol and scorer.

PythonEvaluation & Benchmarks
26
typesafe-ai/WorkflowEvals

evals.typesafe.ai workflow code published

PythonEvaluation & Benchmarks
27
hacan359/xerabora

A desktop RetroAchievements client: your library, what to chase next, missable warnings, leaderboards and rank in one window. Plug in a console that streams its memory and the same window goes live. The first console that does is a real PlayStation 2, through a patched Open PS2 Loader.

CEvaluation & BenchmarksSelf-Hosted Apps