Tag · 2 repos
llm-evaluation
Repositories carrying the llm-evaluation tag.
Exact tag match2 sample repositories
4
127
ifixai-ai/iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
Pythonagent-evaluationaiai-alignment
18.4K
1.4K forks
+40
today
PyModel/jev-judge-mcp
Typed judgment tools for MCP agents. TypeSafe's Jev model as verify, screen, find, classify, rerank, decide, compare, extract, review, gate, and score: the model judges, policy decides auto, review, or escalate.
Pythonagent-toolsai-agentsai-guardrails
84
4 forks
+3
today