Tag · 6 repos
evals
Repositories carrying the evals tag.
The platform for continuously improving AI agents.
Framework for evaluating and improving agents
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, a
Vivaria is METR's tool for running evaluations and conducting agent elicitation research.
Methodology and deterministic evaluation tools for iterative improvement of skills, repositories and agent workflows, with independent review and regression checks.
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.