Skip to main content

Tag · 6 repos

evals

Repositories carrying the evals tag.

Exact tag match6 sample repositories
239
overmind-core/overmind

The platform for continuously improving AI agents.

PythonAI AgentsFine-tuning & Training
516
harbor-framework/harbor

Framework for evaluating and improving agents

PythonEvaluation & Benchmarks
536
calmrocks/ai-engineer-notebooks

Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, a

Jupyter NotebookFine-tuning & TrainingLLMOps & Gateways
716
METR/vivaria

Vivaria is METR's tool for running evaluations and conducting agent elicitation research.

TypeScriptEvaluation & Benchmarks
2071
DaizeDong/self-evolve

Methodology and deterministic evaluation tools for iterative improvement of skills, repositories and agent workflows, with independent review and regression checks.

PythonAI AgentsEvaluation & Benchmarks
2078
FailproofAI/failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.

TypeScriptAI AgentsAI Coding Assistants