[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.
Top AI-EVALUATION GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #ai-evaluation.
[Public preview] The first public quant trading benchmark scored against what professional traders actually made on the same asset
[Public preview] An advanced AI Benchmark for math, produced by mathmo at St John's College, Cambridge.
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
A reproducible local LLM benchmarking platform with versioned datasets, deterministic scoring, and auditable reports.