Claude Sonnet 5.5 Benchmarks: Comparison with Opus and other models
Anthropic has officially announced Claude Sonnet 5.5, which is a significant upgrade over Claude Sonnet 5 and offers near Opus 5.5 level performance. In this post, here is a full benchmark comparison of Claude Sonnet 5.5 with Claude Sonnet 5 and also with Opus 5.5 in terms of performance, AI benchmarks, pricing, and much more.
Benchmarks: Sonnet 5 vs Sonnet 5.5
| Benchmark | Claude Sonnet 5 | Claude Sonnet 5.5 | Change |
|---|---|---|---|
| Terminal-Bench 4.0 | 10.3% | 70.6% | +60.3 points |
| FrontierCode 1.1 (Main) | 42.4% | 52.1%Xhigh effort. 46.2% at Max | +9.7 points |
| CursorBench 4.0 | 34.1% | 55.5% | +21.4 points |
| GDPval-AA v2.1 | 1449 | 1844 | +395 |
| AA-Briefcase v1.1 | 1359 | 1811 | +452 |
| Humanity's Last Exam | 54.9% | 64.5% | +9.6 points |
| OSWorld 2.1 | 57.0% | 80.1% | +23.1 points |
| Chartography | 15.6% | 61.6% | +46.0 points |
Sonnet 5.5 vs Opus 5.5
| Benchmark | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% |
| FrontierCode 1.1 (Main) | 52.1%Xhigh effort | 54.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1846 |
| AA-Briefcase v1.1 | 1811 | 1822 |
| Humanity's Last Exam | 64.5% | 67.7% |
| OSWorld 2.1 | 80.1% | 81.8% |
| Chartography | 61.6% | 64.4% |
Sonnet 5.5 vs other models
| Benchmark | Claude Sonnet 5.5 | Claude Opus 5.5 | GPT-6 Astra | GPT-6 Sol | Grok 4.7 | Gemini 3.8 Flash | DeepSeek V4.1 Flash | DeepSeek V4 Pro | Kimi K3 | Claude Fable 5.1 | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | 57.9% | - | 37.6% | 19.1% | 31.2% | - | - | 55.8% | 37.3% |
| FrontierCode 1.1 (Main) | 52.1%Xhigh effort | 54.4% | 53.3% | 49.3% | - | 43.6% | - | - | - | 50.3% | 47.5% |
| CursorBench 4.0 | 55.5% | 57.8% | - | - | 46.3% | - | - | - | - | 51.8% | 41.7% |
| GDPval-AA v2.1 | 1844 | 1846 | 1542 | 1487 | - | - | - | - | - | 1735 | 1588 |
| AA-Briefcase v1.1 | 1811 | 1822 | - | 1483 | 1657 | - | - | - | - | 1678 | 1487 |
| AutomationBench | - | 40.0% | 41.4% | - | - | - | 54.8% | 31.8% | - | 31.4% | 28.8% |
| DeepSWE v1.1 | - | - | 74.1% | - | 71.0% | 73.7% | 74.2% | 62.7% | 67.5% | 67.4% | 72.7% |
| Terminal-Bench 2.1 | - | - | - | - | - | 89.4% | 90.6% | 87.9% | 88.3% | - | 88.8% |
| Humanity's Last Exam | 64.5% | 67.7% | 57.2% | - | - | - | 63.9% | 60.0% | - | 65.6% | - |
| Terminal-Bench Science 0.1 | - | 58.7% | 64.6% | - | - | - | - | - | - | 52.6% | 22.4% |
| OSWorld 2.1 | 80.1% | 81.8% | - | - | - | - | - | - | - | 80.7% | - |
| Chartography, no tools | 61.6% | 64.4% | - | 53.6% | - | - | - | - | - | - | - |
| Chartography, with tools | - | 89.0% | - | - | - | - | 78.9% | - | - | 88.4% | - |
| GPQA Diamond | - | - | 96.0% | - | - | 95.3% | 90.9% | - | - | 93.7% | 94.6% |
| FrontierMath Tier 4 v2 | - | - | 97.6% | - | - | - | - | - | - | 87.8% | 83.0% |
| HealthBench Professional | - | - | 63.4% | - | - | 52.1% | - | - | - | 58.1% | 60.5% |
| Vals Finance Agent v2 | - | - | - | - | - | 61.4% | - | - | - | - | 53.8% |
| Harvey Legal Agent | - | - | - | - | 19.6% | 10.0% | - | - | - | - | - |
A dash means the lab has not reported a score for that model.
Pricing and specs
| Spec | Claude Sonnet 5 | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|---|
| Input | $2 | $2 | $4 |
| Output | $10 | $10 | $20 |
| Cache reads | $0.20 | $0.20 | $0.20 |
| Cache writes | $2.50 | $2.50 | $5 |
| Batch API | 50% off | 50% off | 50% off |
| Cost per task | Baseline | Up to 30% lessFewer tokens for the same work | HigherTwice the token price |
| Output speed | Baseline | 30%+ fasterFastest Sonnet yet | SlowerModerate latency |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens |
| Thinking | Adaptive | Adaptive | AdaptiveAlways on |
| Default effort | High | High | Medium |
| Knowledge cutoff | January 2026 | June 2026 | June 2026 |
| Released | June 30, 2026 | September 28, 2026 | September 22, 2026 |
| API model ID | claude-sonnet-5 | claude-sonnet-5-5 | claude-opus-5-5 |
Sonnet 5.5 keeps Sonnet 5's price but needs fewer tokens per task, so the same job usually costs less.
API pricing vs other models
| Model | Input, per 1M tokens | Output, per 1M tokens |
|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Gemini 3.7 Flash | $0.75 | $3.75 |
| Grok 4.7 | $2 | $6 |
| Gemini 3.5 Flash | $1.50 | $9 |
| Claude Sonnet 5.5 | $2 | $10 |
| Claude Sonnet 5 | $2 | $10 |
| GPT-6 Sol | $2 | $10 |
| GPT-5.6 Terra | $2 | $12 |
| Claude Opus 5.5 | $4 | $20 |
| GPT-5.6 Sol | $4 | $20 |
| Claude Opus 5 | $5 | $25 |
| GPT-6 Astra | $10 | $50 |
| Claude Fable 5.1 | $10 | $50 |