Claude Haiku 5.5 vs GPT-6 Luna: Which One Is Better?
Anthropic has officially announced Claude Haiku 5.5, its fastest and cheapest model ever, to take on GPT-6 Luna. Here is how they compare with each other.
Benchmarks
| Benchmark | Claude Haiku 5.5 | GPT-6 Luna | Difference |
|---|---|---|---|
| GDPval-AA v2.1 | 1620 Elo | 1437 Elo | Haiku +183 Elo |
| AA-Briefcase v1.1 | 1578 Elo | 1336 Elo | Haiku +242 Elo |
| OSWorld 2.1 (offline subset) | 72.4% | 48.9% | Haiku +23.5 points |
| Terminal-Bench 4.0 | 39.2% | 16.4% | Haiku +22.8 points |
| FrontierCode 1.1 (Main) | 46.4% | 42.4% | Haiku +4.0 points |
| Chartography, no tools | 46.4% | 29.1% | Haiku +17.3 points |
Artificial Analysis benchmarks
| Benchmark | Claude Haiku 5.5 (Max) | GPT-6 Luna (Max) | Difference |
|---|---|---|---|
| Intelligence Index | 43 | 38 | Haiku +5 |
| AA-Briefcase v1.1 | 1578 Elo | 1336 Elo | Haiku +242 Elo |
| GDPval-AA v2.1 | 1620 Elo | 1437 Elo | Haiku +183 Elo |
| AutomationBench-AA | 35% | 53% | Luna +18 points |
| Terminal-Bench 4.0 | 33% | 13% | Haiku +20 points |
| SciCode | 55% | 55% | Tie |
| Humanity's Last Exam | 44% | 39% | Haiku +5 points |
| GDP.pdf | 21% | 23% | Luna +2 points |
| CritPt | 19% | 19% | Tie |
| AA-Omniscience | 11 | 1 | Haiku +10 |
| AA-LCR v1.1 | 83% | 83% | Tie |
Cost and efficiency
| Metric | Claude Haiku 5.5 (Max) | GPT-6 Luna (Max) | Difference |
|---|---|---|---|
| Intelligence Index | 43 | 38 | Haiku +5 |
| Cost per task | $0.21 | $0.07 | Luna about 67% cheaper |
| Cost to run the Intelligence Index | $330 | $122 | Luna about 63% cheaper |
| Output speed | 243 tokens/s | 128 tokens/s | Haiku about 1.9x faster |
| Input, per 1M tokens | $0.10 | $0.10 | Same |
| Cached input, per 1M tokens | $0.01 | $0.01 | Same |
| Output, per 1M tokens | $0.50 | $0.50 | Same |
Pricing and specs
| Specification | Claude Haiku 5.5 | GPT-6 Luna | Difference |
|---|---|---|---|
| Input, prompts up to 100K | $0.10 | $0.10 | Same |
| Cached input, prompts up to 100K | $0.01 | $0.01 | Same |
| Cache write, prompts up to 100K | $0.125 | $0.125 | Same |
| Output, prompts up to 100K | $0.50 | $0.50 | Same |
| Input, prompts from 100K to 272K | $0.50 | $0.10 | Luna much cheaper |
| Output, prompts from 100K to 272K | $2.50 | $0.50 | Luna much cheaper |
| Input, prompts over 272K | $0.50 | $0.20 | Luna 60% cheaper |
| Output, prompts over 272K | $2.50 | $0.75 | Luna 70% cheaper |
| Context window | 1M tokens | 1.05M tokens | Nearly identical |
| Max output | 128K tokens | 128K tokens | Same |
| Reasoning effort | Low to Max | None to Max | Luna can also skip reasoning |
| Text input | Yes | Yes | Same |
| Image input | Yes | Yes | Same |