Bounded, evidence-driven SKILL.md evolution under frozen evaluation and safety contracts.
Top EVALUATION GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #evaluation.
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
$mol - fastest reactive micro-modular compact flexible lazy ui web framework.
Improve your OpenSearch, Elasticsearch, Solr, Vectara, Algolia and Custom Search search quality.
MimIR is my Intermediate Representation
A live arena where LLM agents trade real market data with virtual money — every thought, tool call, thesis and post-mortem is public. Bring your own model and key.
Evaluation and Tracking for LLM Experiments and AI Agents
[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.
[Public preview] The first public quant trading benchmark scored against what professional traders actually made on the same asset
Robot-agent harness with a CLI and local Web console for LIBERO short evaluation
[Public preview] An advanced AI Benchmark for math, produced by mathmo at St John's College, Cambridge.
Agent evals and guardrails as Jev decisions: one request per trace, a fraction of a cent, fast enough for the agent loop. Runs locally with Kev or Laya.
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
AgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, and GEAK—can be evaluated side-by-side on the same GPU kernel tasks, using objective and reproducible metrics.
A reproducible local LLM benchmarking platform with versioned datasets, deterministic scoring, and auditable reports.
LangSmith Client SDK Implementations