Gemini 4 Argon vs Claude Opus 5.5
Google has officially announced Gemini 4 Argon. Here is how it compares with Claude Opus 5.5. Gemini 4 Argon is cheaper than Claude, but Opus 5.5 leads in many benchmarks.
Benchmarks
| Benchmark | Gemini 4 Argon | Claude Opus 5.5 | Difference |
|---|---|---|---|
| Vals Index | 68.9% | 67.0% | Argon +1.9 points |
| AutomationBench | 51.3% | 42.5% | Argon +8.8 points |
| Vals Finance Agent v2 | 65.4% | 58.6% | Argon +6.8 points |
| Harvey Legal Agent | 19.6% | 3.8% | Argon +15.8 points |
| DeepSWE v1.1 | 77.9% | 74.2% | Argon +3.7 points |
| FrontierSWE v2 | 55.0% | 62.3% | Opus +7.3 points |
| Vibe Code Bench | 91.9% | 90.3% | Argon +1.6 points |
| Terminal-Bench 4.0 | 57.4% | 66.4% | Opus +9.0 points |
| PostTrainBench | 45.3% | 49.3% | Opus +4.0 points |
| Terminal-Bench Science 0.1 | 57.6% | 63.3% | Opus +5.7 points |
| LABBench 2 | 88.8% | 73.1% | Argon +15.7 points |
| RiemannBench | 76.0% | 69.6% | Argon +6.4 points |
| GraphWalks up to 128K | 99.7% | 90.6% | Argon +9.1 points |
| GraphWalks 256K to 1M | 84.2% | 66.8% | Argon +17.4 points |
| Agent's Last Exam | 39.5% | 38.2% | Argon +1.3 points |
| OSWorld 2.0 | 69.2% | - | - |
| Chartography | 71.6% | 66.3% | Argon +5.3 points |
| LVBench | 91.7% | 83.7% | Argon +8.0 points |
| CWE-bench v1 | 68.0% | 67.0% | Argon +1.0 point |
Artificial Analysis benchmarks
| Benchmark | Gemini 4 Argon (High) | Claude Opus 5.5 (Max) | Difference |
|---|---|---|---|
| Intelligence Index v4.3.2 | 53 | 58 | Opus +5 |
| AA-Briefcase v1.1 | 1494 Elo | 1822 Elo | Opus +328 Elo |
| GDPval-AA v2.1 | 1611 Elo | 1846 Elo | Opus +235 Elo |
| AutomationBench-AA | 78% | 70% | Argon +8 points |
| Terminal-Bench 4.0 | 57% | 60% | Opus +3 points |
| SciCode | 62% | 67% | Opus +5 points |
| Humanity's Last Exam | 57% | 61% | Opus +4 points |
| GDP.pdf | 22% | 26% | Opus +4 points |
| CritPt | 27% | 32% | Opus +5 points |
| AA-Omniscience | 42 | 46 | Opus +4 |
| AA-LCR v1.1 | 80% | 85% | Opus +5 points |
Cost and efficiency
| Metric | Gemini 4 Argon (High) | Claude Opus 5.5 (Max) | Difference |
|---|---|---|---|
| Intelligence Index | 53 | 58 | Opus +5 |
| Cost per Intelligence Index task | $1.99 | $5.98 | Argon about 67% cheaper |
| Output tokens per task | 62K | 119K | Argon uses about 48% fewer |
| Reasoning tokens per task | 36K | 84K | Argon uses about 57% fewer |
| Output tokens across the Index | 113M | 260M | Argon uses about 57% fewer |
| Input, per 1M tokens | $2 | $4 | Argon 50% cheaper |
| Output, per 1M tokens | $10 | $20 | Argon 50% cheaper |
| Cached input, per 1M tokens | $0.10 | $0.20 | Argon 50% cheaper |
Pricing and specs
| Specification | Gemini 4 Argon | Claude Opus 5.5 | Difference |
|---|---|---|---|
| Input now, per 1M tokens | $2Introductory | $4 | Argon 50% cheaper |
| Output now, per 1M tokens | $10Introductory | $20 | Argon 50% cheaper |
| Cached input now, per 1M tokens | $0.10 | $0.20 | Argon 50% cheaper |
| Input later, per 1M tokens | $4 | $4 | Same |
| Output later, per 1M tokens | $20 | $20 | Same |
| Context window | 1M tokens | 1M tokens | Same |
| Max output | Up to 1M tokens | Not stated at launch | - |
| Reasoning | Yes | Adaptive | Both |
| Text input | Yes | Yes | Same |
| Image input | Yes | Yes | Same |
| Availability | Limited rollout at first | Available to everyone | Opus is wider |