Google Gemini 4 Argon Benchmarks: Is Google Back?
Google has officially announced its latest frontier model, Gemini 4 Argon, but it is not available to everyone at launch. It is a significant upgrade over Gemini 3.1 Pro, and it also beats Gemini 3.8 Flash and other frontier models in certain benchmarks. Here are all the latest Gemini 4 Argon benchmarks.
Benchmarks: Gemini 3.8 Flash vs Gemini 4 Argon
| Benchmark | Gemini 3.8 Flash | Gemini 4 Argon | Change |
|---|---|---|---|
| DeepSWE v1.1 | 73.7% | 77.9% | +4.2 points |
| Finance Agent v2 | 61.4% | 65.4% | +4.0 points |
| Harvey Legal Agent | 10.0% | 19.6% | +9.6 points |
| Terminal-Bench 4.0 | 19.1% | 57.4% | +38.3 points |
| OSWorld 2.0 | 59.0% | 69.2% | +10.2 points |
| LVBench | 87.8% | 91.7% | +3.9 points |
| LABBench 2 | 86.2% | 88.8% | +2.6 points |
| Max output | 64K tokens | 1M tokens | About 15.6x |
Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 67.0% |
| AutomationBench | 51.3% | 41.4% | 42.5% |
| Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| Harvey Legal Agent | 19.6% | 5.4% | 3.8% |
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 66.4% |
| Terminal-Bench Science | 57.6% | 68.1% | 63.3% |
| LABBench 2 | 88.8% | 85.4% | 73.1% |
| GraphWalks 256K to 1M | 84.2% | 71.8% | 66.8% |
| OSWorld 2.0 | 69.2% | 72.6% | - |
| Chartography | 71.6% | 71.0% | 66.3% |
| LVBench | 91.7% | 87.5% | 83.7% |
| CWE-bench v1 | 68.0% | 68.0% | 67.0% |
A dash means no score has been reported for that model.
API pricing
| Model | Input, per 1M tokens | Cached input | Output, per 1M tokens |
|---|---|---|---|
| Gemini 4 Argon | $2 | $0.10 | $10 |
| GPT-6.1 Sol | $2 | $0.10 | $10 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |