Evaluation and Tracking for LLM Experiments and AI Agents
Topic Catalog
Top AGENT-EVALUATION GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #agent-evaluation.
SPONSORED BY
3 matching projects · updated Sep 29, 2026
How we rank ↗
↑ +1 today
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
↑ +1 today
AgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, and GEAK—can be evaluated side-by-side on the same GPU kernel tasks, using objective and reproducible metrics.
↑ +1 today
From our network