Gemini 4 Argon vs Gemini 3.8 Flash

Finally, a pro model from Google Gemini. Google has officially announced Gemini 4 Argon, and here is how it compares with Gemini 3.8 Flash, which was announced recently.

Benchmarks

BenchmarkGemini 4 ArgonGemini 3.8 FlashDifference
DeepSWE v1.177.9%73.7%Argon +4.2 points
Vals Finance Agent v265.4%61.4%Argon +4.0 points
Harvey Legal Agent19.6%10.0%Argon +9.6 points
Terminal-Bench 4.057.4%19.1%Argon +38.3 points
LABBench 288.8%86.2%Argon +2.6 points
OSWorld 2.0 (partial credit)69.2%59.0%Argon +10.2 points
LVBench91.7%87.8%AgenticArgon +3.9 points

Artificial Analysis benchmarks

BenchmarkGemini 4 Argon (High)Gemini 3.8 Flash (High)Difference
Intelligence Index v4.3.25341Argon +12
AA-Briefcase v1.11494 Elo1202 EloArgon +292 Elo
GDPval-AA v2.11611 Elo1412 EloArgon +199 Elo
AutomationBench-AA78%60%Argon +18 points
Terminal-Bench 4.057%20%Argon +37 points
SciCode62%57%Argon +5 points
Humanity's Last Exam57%48%Argon +9 points
GDP.pdf22%21%Argon +1 point
CritPt27%18%Argon +9 points
AA-Omniscience4230Argon +12
AA-LCR v1.180%81%3.8 Flash +1 point

Cost and efficiency

MetricGemini 4 Argon (High)Gemini 3.8 Flash (High)Difference
Intelligence Index5341Argon +12
Cost per Intelligence Index task$1.99$1.243.8 Flash about 38% cheaper
Output tokens per task62K71KArgon uses about 13% fewer
Reasoning tokens per task36K43KArgon uses about 16% fewer
Output tokens across the Index113M172MArgon uses about 34% fewer
Input, per 1M tokens$2Introductory$0.75Introductory3.8 Flash 62.5% cheaper
Output, per 1M tokens$10Introductory$3.75Introductory3.8 Flash 62.5% cheaper
Cached input, per 1M tokens$0.10Introductory$0.075Introductory3.8 Flash 25% cheaper

Pricing and specs

SpecificationGemini 4 ArgonGemini 3.8 FlashDifference
Input now, per 1M tokens$2Introductory$0.75Introductory3.8 Flash about 2.7x cheaper
Cached input now, per 1M tokens$0.10Introductory$0.075Introductory3.8 Flash cheaper
Output now, per 1M tokens$10Introductory$3.75Introductory3.8 Flash about 2.7x cheaper
Input later, per 1M tokens$4Standard$1.50Standard3.8 Flash about 2.7x cheaper
Output later, per 1M tokens$20Standard$7.50Standard3.8 Flash about 2.7x cheaper
Context window1M tokens1M tokensSame
Max output1M tokens64K tokensArgon about 15.6x larger
ThinkingYesYes, Low, Medium and HighBoth
Public API statusLimited rolloutGenerally available3.8 Flash is wider
ReleasedSeptember 30, 2026September 2, 2026Argon is 28 days newer