← Back to all entries
2025-11-29 ✅ Best Practices

What Opus 4.5's Benchmarks Actually Measure — and Why Vending-Bench Matters

What Opus 4.5's Benchmarks Actually Measure — and Why Vending-Bench Matters — visual for 2025-11-29

What Opus 4.5's Benchmarks Actually Measure — and Why Vending-Bench Matters

A week's worth of headlines about Claude Opus 4.5 have leaned on a handful of benchmark names — SWE-bench, Aider Polyglot, Vending-Bench — without always explaining what each one is actually testing. On a quiet Saturday, it's worth slowing down and reading the model card the way you'd read any other evaluation: for what it measures, not just the number attached to it.

Four different things, four different signals

The best-practice takeaway

No single benchmark tells you how a model will behave on your actual workload. A model that's state-of-the-art on SWE-bench but mediocre on Vending-Bench is telling you it's better at fixing an isolated bug than at running a long, stateful process unsupervised. Match the benchmark to the shape of the task you're actually planning to hand it, and treat the rest as context.

⭐⭐⭐ anthropic.com
benchmarks evaluation SWE-bench Vending-Bench retrospective
Source trust ratings ⭐⭐⭐ Official Anthropic  ·  ⭐⭐ Established press  ·  Community / research