Artificial Analysis benchmarks
| Benchmark | GPT-6.1 Sol (Max) | Claude Sonnet 5.5 (Max) | Difference |
|---|
| Intelligence Index v4.3.2 | 51.8 | 56.0 | Sonnet +4.2 |
| AA-Briefcase v1.1 | 1564.2 Elo | 1811 Elo | Sonnet +246.8 Elo |
| GDPval-AA v2.1 | 1575.1 Elo | 1844 Elo | Sonnet +268.9 Elo |
| AutomationBench-AA | 64.9% | 71% | Sonnet +6.1 points |
| Terminal-Bench 4.0 | 56.1% | 64% | Sonnet +7.9 points |
| SciCode | 54.2% | 61% | Sonnet +6.8 points |
| Humanity's Last Exam | 52.9% | 55% | Sonnet +2.1 points |
| GDP.pdf | 31.0% | 26% | GPT-6.1 Sol +5.0 points |
| CritPt | 31.7% | 31% | GPT-6.1 Sol +0.7 points |
| AA-Omniscience | 41.5 | 32 | GPT-6.1 Sol +9.5 |
| AA-LCR v1.1 | 83% | 83% | Tie |
Cost and efficiency
| Metric | GPT-6.1 Sol (Max) | Claude Sonnet 5.5 (Max) | Difference |
|---|
| Intelligence Index | 51.8 | 56.0 | Sonnet +4.2 |
| Cost per Intelligence Index task | $0.72 | $7.60 | Sol about 90.5% cheaper |
| Output tokens per task | About 38K | About 193K | Sol uses about 80% fewer |
| Output speed | About 67 tokens/s | About 138 tokens/s | Sonnet about 2.1x faster |
| Input, per 1M tokens | $2 | $2 | Same |
| Output, per 1M tokens | $10 | $10 | Same |
| Cached input, per 1M tokens | $0.10 | $0.20 | Sol 50% cheaper |
Pricing and specs
| Specification | GPT-6.1 Sol | Claude Sonnet 5.5 | Difference |
|---|
| Input, per 1M tokens | $2 | $2 | Same |
| Cached input, per 1M tokens | $0.10 | $0.20 | Sol 50% cheaper |
| Cache write, per 1M tokens | $2.50 | $2.50 | Same |
| Output, per 1M tokens | $10 | $10 | Same |
| Context window | 1.05M tokens | 1M tokens | Sol +50K tokens |
| Max output | 128K tokens | 128K tokens | Same |
| Reasoning levels | Low to Max | Low to Max, adaptive | Both adjustable |
| Long-context surcharge | Yes, above 272K input tokens | None stated | Different pricing model |