Skip to main content

Tag · 2 repos

llm-evaluation

Repositories carrying the llm-evaluation tag.

Exact tag match2 sample repositories
4
ifixai-ai/iFixAi

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

Pythonagent-evaluationaiai-alignment
127
PyModel/jev-judge-mcp

Typed judgment tools for MCP agents. TypeSafe's Jev model as verify, screen, find, classify, rerank, decide, compare, extract, review, gate, and score: the model judges, policy decides auto, review, or escalate.

Pythonagent-toolsai-agentsai-guardrails