← Home

Benchmarks & Tests

Head-to-head agent tests with methodology, scores, and honest notes. Not sponsored. Not cherry-picked. Just reproducible recipes and what actually happened.

Showing 22 of 22 benchmarks

Coding BenchmarkVerified 1 day ago

SWE-bench Verified Leaderboard — July 2026

Claude Opus 5 takes the top spot at 96–97% on SWE-bench Verified, edging out GPT-5.6 Sol and Claude Fable 5 as the frontier coding benchmark tightens at the top.

2026-07-29·7 agents
Winner:Claude Opus 5
codingbenchmarksgithubagents
View results
Agent Safety BenchmarkVerified 8 days ago

StepShield: Temporal Agent Guardrails Benchmark

NeurIPS 2026 benchmark with 9,429 trajectories and step-level annotations for evaluating when (not whether) to intervene on rogue AI agents — temporal evaluation of guardrails.

2026-07-22·1 agents
Winner:Tie
agent-safetyguardrailsbenchmarktemporal-evaluation
View results
Tool-Use BenchmarkVerified 17 days ago

tool-eval-bench: Tool-Calling Quality Across Serving Stacks

69 deterministic tool-use scenarios (plus optional Hard Mode) for OpenAI-compatible endpoints — selection, parameters, chains, restraint, recovery, safety, and structured output.

2026-07-13·1 agents
Winner:Tie
tool-callingbenchmarksagentsvllm
View results
Tool-Use BenchmarkVerified 17 days ago

ToolCall-15: Deterministic Tool-Use Bench Pack

Fifteen scenarios across selection, parameter precision, multi-step chains, restraint, and error recovery — a reproducible tool-use pack for BenchLocal and CLI runners.

2026-07-13·1 agents
Winner:Tie
tool-callingbenchmarksagentsbenchlocal
View results
Integration BenchmarkVerified 29 days ago

Multi-Turn Tool Consistency Benchmark

Measures whether agents keep track of prior tool outputs and use them correctly across several turns without losing the thread.

2026-07-01·4 agents
Winner:Claude Code
memorytool-callingcontextagents
View results
Integration BenchmarkVerified 29 days ago

Agentic Web Navigation Benchmark

Tests how well agents can navigate real websites: find information, fill forms, click buttons, and recover from common UI changes.

2026-07-01·4 agents
Winner:Manus
webbrowseragentsbenchmark
View results
Coding BenchmarkVerified 29 days ago

Local LLM Reasoning Benchmark

Compares open-weight local models on reasoning, math, and code tasks to find the best on-premise balance of capability and cost.

2026-07-01·4 agents
Winner:Qwen3.6-27B
local-modelsreasoningcodingbenchmark
View results
Agent BenchmarkVerified 36 days ago

Long-Context RAG Recall Benchmark

New benchmark measuring how well agents and models retrieve facts from very long documents. Long-context models are closing the gap with vector RAG on many tasks.

2026-06-24·4 agents
Winner:Gemini 2.5 Pro
agentsbenchmarksraglong-context
View results
Agent BenchmarkVerified 39 days ago

Agents' Last Exam (ALE) — June 2026 Launch

UC Berkeley's new long-horizon professional-work benchmark shows even top agents failing most tasks. GPT-5.5/Codex leads at 24.0%, Claude Fable 5 at 22.0%.

2026-06-21·4 agents
Winner:Codex (GPT-5.5)
agentsbenchmarkslong-horizonprofession-tasks
View results
Coding BenchmarkVerified 39 days ago

SWE-bench Verified Leaderboard — June 2026

Real-world coding benchmark update: Claude Mythos 5 takes the top spot at ~78%, DeepSeek V4.1 Pro sits within 6 points of frontier closed models.

2026-06-21·6 agents
Winner:Claude Mythos 5
codingbenchmarksgithubagents
View results
Code-Quality BenchmarkVerified 39 days ago

Sonar GPT-5.5 Java Evaluation — June 2026

Independent 4,444-task Java benchmark: GPT-5.5 shows one of the cleanest security profiles Sonar has measured, but concurrency bugs and verification debt are real risks.

2026-06-21·1 agents
Winner:N/A — single-model evaluation
code-qualitysecurityjavagpt-5.5
View results
Coding BenchmarkVerified 44 days ago

API Error Handling Benchmark

Which agents produce correct, defensive error handling for REST endpoints with edge cases?

2026-06-16·4 agents
Winner:Aider
apierror-handlingagentsbenchmark
View results
Coding BenchmarkVerified 44 days ago

Local Model Code Generation Benchmark

Can local models keep up with cloud APIs on everyday coding tasks? We tested Qwen, Llama, and Mistral against Claude.

2026-06-16·4 agents
Winner:Claude 3.5 Haiku via API
local-modelscodingbenchmarkcost
View results
No-Code BenchmarkVerified 44 days ago

Agent UI Generation Benchmark

Which agents turn a prompt into a clean, responsive UI fastest? React + Tailwind dashboard edition.

2026-06-16·4 agents
Winner:v0
uireacttailwindno-code
View results
Integration BenchmarkVerified 44 days ago

Agent Memory Benchmark

Which agents remember context across a multi-turn conversation and use it correctly in later tasks?

2026-06-16·4 agents
Winner:Letta
memorycontextagentsbenchmark
View results
Security BenchmarkVerified 44 days ago

Prompt Injection Resilience Benchmark

Tested leading agents against direct, indirect, and role-play injection attacks. See who stayed on task.

2026-06-16·4 agents
Winner:Claude 3.7 Sonnet
securityprompt-injectionagentsbenchmark
View results
Integration BenchmarkVerified 45 days ago

MCP Tool Calling Benchmark

Compared agents on correctly invoking Model Context Protocol servers: filesystem, GitHub, and Postgres.

2026-06-15·4 agents
Winner:Claude Code
mcpmodel-context-protocoltoolsagents
View results
Coding BenchmarkVerified 45 days ago

Local Model Coding Benchmark

Compared quantized local models running through Ollama on real coding tasks: refactoring, bug fixing, and test generation.

2026-06-14·4 agents
Winner:qwen2.5-coder:14b
local-modelsollamaquantizationcoding
View results
Security BenchmarkVerified 45 days ago

Agent Security Audit Benchmark

Tested whether agents can find prompt injection vectors, unsafe tool permissions, and secret leakage in a deliberately vulnerable agent app.

2026-06-12·4 agents
Winner:Claude Code
securityprompt-injectionauditagents
View results
No-Code BenchmarkVerified 45 days ago

Prompt-to-App Benchmark

Measured how far four no-code agents get from a single prompt: a task tracker with auth and a database.

2026-06-10·4 agents
Winner:Lovable
no-codevibe-codingauthdatabase
View results
Coding BenchmarkVerified 45 days ago

UI Regression Fix Benchmark

Compared agents on finding and fixing a responsive CSS regression in an existing React app.

2026-06-08·4 agents
Winner:Cursor
cssreactregressionui
View results
Coding BenchmarkVerified 45 days ago

Next.js CRUD App Benchmark

Measured how long each agent takes to scaffold a working Next.js CRUD app with authentication.

2026-06-01·4 agents
Winner:Claude Code
nextjscrudauthagents
View results