Bounded, evidence-driven SKILL.md evolution under frozen evaluation and safety contracts.
Top LLM-EVALUATION GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #llm-evaluation.
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
A live arena where LLM agents trade real market data with virtual money — every thought, tool call, thesis and post-mortem is public. Bring your own model and key.
Evaluation and Tracking for LLM Experiments and AI Agents
Agent evals and guardrails as Jev decisions: one request per trace, a fraction of a cent, fast enough for the agent loop. Runs locally with Kev or Laya.
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
A reproducible local LLM benchmarking platform with versioned datasets, deterministic scoring, and auditable reports.