Cursor's CursorBench 3.1 crowns Fable 5 Max, but Cursor's own model wins on cost
Cursor has published results from CursorBench 3.1, the latest version of its internal coding-agent benchmark, with Anthropic's Fable 5 Max topping the leaderboard and a cheaper Cursor-built model close behind on a cost-adjusted basis.
What's new
Cursor's evals page describes the benchmark's methodology directly: "We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better." Version 3.1 refreshes the problem set, having "introduced problems focused on codebase understanding, bugfinding, planning, and code review" — a shift toward messier, more realistic engineering tasks rather than isolated coding puzzles.
On the new problem set, Fable 5 Max posted the top score, reaching 72.9%, though at a relatively high average cost of $18.02 per task. Cursor's own Composer 2.5 model placed as the strongest low-cost option, scoring 63.2% while costing just $0.55 per task — roughly 3% of Fable 5 Max's per-task cost for a good chunk of its capability.
Context
CursorBench is Cursor's own internally maintained evaluation, built from real, ambiguous multi-file tasks drawn from actual Cursor user sessions rather than a static public dataset — a design choice meant to track how coding agents perform in the kind of messy, underspecified work that makes up daily engineering, rather than curated leetcode-style problems. Cursor has iterated on the benchmark across multiple versions this year as it pushes both its own in-house models, like the Composer line, and third-party frontier models available inside its editor.
The result also lands amid an active price-versus-capability debate across the coding-agent market, with labs and tool builders increasingly reporting cost-per-task alongside raw accuracy, as cheaper, specialized models close the gap with expensive frontier options on a growing share of everyday engineering tasks.
Why it matters
The results reinforce a split that has become common across coding-agent benchmarks this year: the most capable frontier model does not necessarily win on cost-efficiency. Fable 5 Max's benchmark lead comes at more than 30 times the per-task cost of Composer 2.5, meaning the practical choice for a team running these agents at scale is unlikely to default to the top-scoring model for every task.
For Cursor specifically, having its own Composer 2.5 model post the strongest capability-per-dollar result on Cursor's own benchmark is a notable data point for the company's push to route more of its coding-agent workload to in-house models rather than paying frontier-lab API prices for every request — while still offering Fable 5 Max as the ceiling option for the hardest, most ambiguous tasks.
Corroborating sources
- Cursor
https://cursor.com/evals
“We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.”