← Home

Benchmarks & Tests

Head-to-head agent tests with methodology, scores, and honest notes. Not sponsored. Not cherry-picked. Just reproducible recipes and what actually happened.

Showing 28 of 28 benchmarks

Agent Construction BenchmarkVerified 4 days ago

ττ-bench: Can Coding Agents Build a Customer-Service Agent? (23.9% vs 82.2% Ceiling)

A new arXiv benchmark makes agent construction the task — a coding agent must deliver a complete customer-service agent from real business records, a client, a production API, and a cost budget. The best system passes just 23.9% vs an 82.2% expert ceiling.

2026-09-04·2 agents
Winner:Claude Opus 5 (under Claude Code) — 23.9%
benchmarkagentscoding-agentcustomer-service
View results
Coding BenchmarkVerified 11 days ago

Terminal-Bench v2.1: Claude Fable 5.1 Leads at 91.4%

The verified refresh of Terminal-Bench fixes 28 of 89 tasks and introduces continuous validation. Claude Fable 5.1 takes the top spot at 91.4%, with GPT-5.6 Sol and Claude Opus 5 close behind.

2026-09-02·8 agents
Winner:Claude Fable 5.1
codingbenchmarksagentsterminal
View results
End-to-End BenchmarkVerified 18 days ago

StartupBench: Market-Validated Agent Workflows

A 97-task benchmark grounded in real AI startup products — not researcher-selected tasks. Even the strongest model completes only 30% under strict acceptance. Specialized agents beat general-purpose ones by 11+ points.

2026-08-24·6 agents
Winner:Kimi-K3 (highest avg score); GPT-5.6-sol (highest completion rate)
benchmarkagentse2emarket-validated
View results
Coding Agent BenchmarkVerified 25 days ago

GLM-5.3 vs Frontier Coding Agents: Terminal-Bench 3.0 Showdown

GLM-5.3's post-training gains push Terminal-Bench 3.0 from 4.6 to 28.3 — but how does it compare to Fable 5 and GPT-5.6 Sol on independent benchmarks?

2026-08-19·3 agents
Winner:Claude Fable 5
benchmarkcodingagentsterminal-bench
View results
Agentic BenchmarkVerified 32 days ago

AA-AnalystAgent: Quantitative Analysis on Real Spreadsheets

Artificial Analysis's agentic benchmark testing models on 80 private quantitative analysis questions across 14 business and science domains — scored with pass^5 reliability.

2026-08-10·4 agents
Winner:Claude Opus 5
benchmarkagentsquantitative-analysisreliability
View results
Reliability BenchmarkVerified 39 days ago

Kimi K3 Hallucination Paradox — August 2026

Kimi K3 ranked third on the AI Intelligence Index while its hallucination rate hit 51%. This test examines what happens when a benchmark rewards attempting more questions over getting them right.

2026-08-05·4 agents
Winner:Claude Opus 5
benchmarkagentshallucinationreliability
View results
Coding BenchmarkVerified 46 days ago

SWE-bench Verified Leaderboard — July 2026

Claude Opus 5 takes the top spot at 96–97% on SWE-bench Verified, edging out GPT-5.6 Sol and Claude Fable 5 as the frontier coding benchmark tightens at the top.

2026-07-29·7 agents
Winner:Claude Opus 5
codingbenchmarksgithubagents
View results
Agent Safety BenchmarkVerified 53 days ago

StepShield: Temporal Agent Guardrails Benchmark

NeurIPS 2026 benchmark with 9,429 trajectories and step-level annotations for evaluating when (not whether) to intervene on rogue AI agents — temporal evaluation of guardrails.

2026-07-22·1 agents
Winner:Tie
agent-safetyguardrailsbenchmarktemporal-evaluation
View results
Tool-Use BenchmarkVerified 62 days ago

tool-eval-bench: Tool-Calling Quality Across Serving Stacks

69 deterministic tool-use scenarios (plus optional Hard Mode) for OpenAI-compatible endpoints — selection, parameters, chains, restraint, recovery, safety, and structured output.

2026-07-13·1 agents
Winner:Tie
tool-callingbenchmarksagentsvllm
View results
Tool-Use BenchmarkVerified 62 days ago

ToolCall-15: Deterministic Tool-Use Bench Pack

Fifteen scenarios across selection, parameter precision, multi-step chains, restraint, and error recovery — a reproducible tool-use pack for BenchLocal and CLI runners.

2026-07-13·1 agents
Winner:Tie
tool-callingbenchmarksagentsbenchlocal
View results
Integration BenchmarkVerified 74 days ago

Multi-Turn Tool Consistency Benchmark

Measures whether agents keep track of prior tool outputs and use them correctly across several turns without losing the thread.

2026-07-01·4 agents
Winner:Claude Code
memorytool-callingcontextagents
View results
Integration BenchmarkVerified 74 days ago

Agentic Web Navigation Benchmark

Tests how well agents can navigate real websites: find information, fill forms, click buttons, and recover from common UI changes.

2026-07-01·4 agents
Winner:Manus
webbrowseragentsbenchmark
View results
Coding BenchmarkVerified 74 days ago

Local LLM Reasoning Benchmark

Compares open-weight local models on reasoning, math, and code tasks to find the best on-premise balance of capability and cost.

2026-07-01·4 agents
Winner:Qwen3.6-27B
local-modelsreasoningcodingbenchmark
View results
Agent BenchmarkVerified 81 days ago

Long-Context RAG Recall Benchmark

New benchmark measuring how well agents and models retrieve facts from very long documents. Long-context models are closing the gap with vector RAG on many tasks.

2026-06-24·4 agents
Winner:Gemini 2.5 Pro
agentsbenchmarksraglong-context
View results
Agent BenchmarkVerified 84 days ago

Agents' Last Exam (ALE) — June 2026 Launch

UC Berkeley's new long-horizon professional-work benchmark shows even top agents failing most tasks. GPT-5.5/Codex leads at 24.0%, Claude Fable 5 at 22.0%.

2026-06-21·4 agents
Winner:Codex (GPT-5.5)
agentsbenchmarkslong-horizonprofession-tasks
View results
Coding BenchmarkVerified 84 days ago

SWE-bench Verified Leaderboard — June 2026

Real-world coding benchmark update: Claude Mythos 5 takes the top spot at ~78%, DeepSeek V4.1 Pro sits within 6 points of frontier closed models.

2026-06-21·6 agents
Winner:Claude Mythos 5
codingbenchmarksgithubagents
View results
Code-Quality BenchmarkVerified 84 days ago

Sonar GPT-5.5 Java Evaluation — June 2026

Independent 4,444-task Java benchmark: GPT-5.5 shows one of the cleanest security profiles Sonar has measured, but concurrency bugs and verification debt are real risks.

2026-06-21·1 agents
Winner:N/A — single-model evaluation
code-qualitysecurityjavagpt-5.5
View results
Coding BenchmarkVerified 89 days ago

API Error Handling Benchmark

Which agents produce correct, defensive error handling for REST endpoints with edge cases?

2026-06-16·4 agents
Winner:Aider
apierror-handlingagentsbenchmark
View results
Coding BenchmarkVerified 89 days ago

Local Model Code Generation Benchmark

Can local models keep up with cloud APIs on everyday coding tasks? We tested Qwen, Llama, and Mistral against Claude.

2026-06-16·4 agents
Winner:Claude 3.5 Haiku via API
local-modelscodingbenchmarkcost
View results
No-Code BenchmarkVerified 89 days ago

Agent UI Generation Benchmark

Which agents turn a prompt into a clean, responsive UI fastest? React + Tailwind dashboard edition.

2026-06-16·4 agents
Winner:v0
uireacttailwindno-code
View results
Integration BenchmarkVerified 89 days ago

Agent Memory Benchmark

Which agents remember context across a multi-turn conversation and use it correctly in later tasks?

2026-06-16·4 agents
Winner:Letta
memorycontextagentsbenchmark
View results
Security BenchmarkVerified 89 days ago

Prompt Injection Resilience Benchmark

Tested leading agents against direct, indirect, and role-play injection attacks. See who stayed on task.

2026-06-16·4 agents
Winner:Claude 3.7 Sonnet
securityprompt-injectionagentsbenchmark
View results
Integration BenchmarkVerified 90 days ago

MCP Tool Calling Benchmark

Compared agents on correctly invoking Model Context Protocol servers: filesystem, GitHub, and Postgres.

2026-06-15·4 agents
Winner:Claude Code
mcpmodel-context-protocoltoolsagents
View results
Coding BenchmarkVerified 90 days ago

Local Model Coding Benchmark

Compared quantized local models running through Ollama on real coding tasks: refactoring, bug fixing, and test generation.

2026-06-14·4 agents
Winner:qwen2.5-coder:14b
local-modelsollamaquantizationcoding
View results
Security BenchmarkVerified 90 days ago

Agent Security Audit Benchmark

Tested whether agents can find prompt injection vectors, unsafe tool permissions, and secret leakage in a deliberately vulnerable agent app.

2026-06-12·4 agents
Winner:Claude Code
securityprompt-injectionauditagents
View results
No-Code BenchmarkVerified 90 days ago

Prompt-to-App Benchmark

Measured how far four no-code agents get from a single prompt: a task tracker with auth and a database.

2026-06-10·4 agents
Winner:Lovable
no-codevibe-codingauthdatabase
View results
Coding BenchmarkVerified 90 days ago

UI Regression Fix Benchmark

Compared agents on finding and fixing a responsive CSS regression in an existing React app.

2026-06-08·4 agents
Winner:Cursor
cssreactregressionui
View results
Coding BenchmarkVerified 90 days ago

Next.js CRUD App Benchmark

Measured how long each agent takes to scaffold a working Next.js CRUD app with authentication.

2026-06-01·4 agents
Winner:Claude Code
nextjscrudauthagents
View results