Claude Sonnet 5.5 Benchmarks: Comparison with Opus and other models

Anthropic has officially announced Claude Sonnet 5.5, which is a significant upgrade over Claude Sonnet 5 and offers near Opus 5.5 level performance. In this post, here is a full benchmark comparison of Claude Sonnet 5.5 with Claude Sonnet 5 and also with Opus 5.5 in terms of performance, AI benchmarks, pricing, and much more.

Benchmarks: Sonnet 5 vs Sonnet 5.5

BenchmarkClaude Sonnet 5Claude Sonnet 5.5Change
Terminal-Bench 4.010.3%70.6%+60.3 points
FrontierCode 1.1 (Main)42.4%52.1%Xhigh effort. 46.2% at Max+9.7 points
CursorBench 4.034.1%55.5%+21.4 points
GDPval-AA v2.114491844+395
AA-Briefcase v1.113591811+452
Humanity's Last Exam54.9%64.5%+9.6 points
OSWorld 2.157.0%80.1%+23.1 points
Chartography15.6%61.6%+46.0 points

Sonnet 5.5 vs Opus 5.5

BenchmarkClaude Sonnet 5.5Claude Opus 5.5
Terminal-Bench 4.070.6%66.4%
FrontierCode 1.1 (Main)52.1%Xhigh effort54.4%
CursorBench 4.055.5%57.8%
GDPval-AA v2.118441846
AA-Briefcase v1.118111822
Humanity's Last Exam64.5%67.7%
OSWorld 2.180.1%81.8%
Chartography61.6%64.4%

Sonnet 5.5 vs other models

BenchmarkClaude Sonnet 5.5Claude Opus 5.5GPT-6 AstraGPT-6 SolGrok 4.7Gemini 3.8 FlashDeepSeek V4.1 FlashDeepSeek V4 ProKimi K3Claude Fable 5.1GPT-5.6 Sol
Terminal-Bench 4.070.6%66.4%57.9%-37.6%19.1%31.2%--55.8%37.3%
FrontierCode 1.1 (Main)52.1%Xhigh effort54.4%53.3%49.3%-43.6%---50.3%47.5%
CursorBench 4.055.5%57.8%--46.3%----51.8%41.7%
GDPval-AA v2.11844184615421487-----17351588
AA-Briefcase v1.118111822-14831657----16781487
AutomationBench-40.0%41.4%---54.8%31.8%-31.4%28.8%
DeepSWE v1.1--74.1%-71.0%73.7%74.2%62.7%67.5%67.4%72.7%
Terminal-Bench 2.1-----89.4%90.6%87.9%88.3%-88.8%
Humanity's Last Exam64.5%67.7%57.2%---63.9%60.0%-65.6%-
Terminal-Bench Science 0.1-58.7%64.6%------52.6%22.4%
OSWorld 2.180.1%81.8%-------80.7%-
Chartography, no tools61.6%64.4%-53.6%-------
Chartography, with tools-89.0%----78.9%--88.4%-
GPQA Diamond--96.0%--95.3%90.9%--93.7%94.6%
FrontierMath Tier 4 v2--97.6%------87.8%83.0%
HealthBench Professional--63.4%--52.1%---58.1%60.5%
Vals Finance Agent v2-----61.4%----53.8%
Harvey Legal Agent----19.6%10.0%-----

A dash means the lab has not reported a score for that model.

Pricing and specs

SpecClaude Sonnet 5Claude Sonnet 5.5Claude Opus 5.5
Input$2$2$4
Output$10$10$20
Cache reads$0.20$0.20$0.20
Cache writes$2.50$2.50$5
Batch API50% off50% off50% off
Cost per taskBaselineUp to 30% lessFewer tokens for the same workHigherTwice the token price
Output speedBaseline30%+ fasterFastest Sonnet yetSlowerModerate latency
Context window1M tokens1M tokens1M tokens
Max output128K tokens128K tokens128K tokens
ThinkingAdaptiveAdaptiveAdaptiveAlways on
Default effortHighHighMedium
Knowledge cutoffJanuary 2026June 2026June 2026
ReleasedJune 30, 2026September 28, 2026September 22, 2026
API model IDclaude-sonnet-5claude-sonnet-5-5claude-opus-5-5

Sonnet 5.5 keeps Sonnet 5's price but needs fewer tokens per task, so the same job usually costs less.

API pricing vs other models

ModelInput, per 1M tokensOutput, per 1M tokens
GPT-6 Luna$0.10$0.50
GPT-5.6 Luna$0.20$1.20
Gemini 3.5 Flash-Lite$0.30$2.50
Gemini 3.8 Flash$0.75$3.75
Gemini 3.7 Flash$0.75$3.75
Grok 4.7$2$6
Gemini 3.5 Flash$1.50$9
Claude Sonnet 5.5$2$10
Claude Sonnet 5$2$10
GPT-6 Sol$2$10
GPT-5.6 Terra$2$12
Claude Opus 5.5$4$20
GPT-5.6 Sol$4$20
Claude Opus 5$5$25
GPT-6 Astra$10$50
Claude Fable 5.1$10$50

Real-world outputs