OpenAI's GPT-5.6 Sol beats Claude Fable 5 by 13.1 points on Agents' Last Exam and costs a third less on the Artificial Analysis Coding Agent Index. That's the headline OpenAI wants you to read. It's also only half the scoreboard. On Artificial Analysis's broader intelligence index, independent measurements still put Fable 5 ahead of Sol, and the gap OpenAI is celebrating exists mainly on the two benchmarks it chose to lead with.

GPT-5.6 launched broadly on July 9, 2026, after a 12-day government-gated preview limited to roughly 20 trusted partners. The family ships in three tiers: Sol, Terra, and Luna, each priced to undercut Fable 5 by a wide margin. Terra costs $2.50 per million input tokens and $15 per million output tokens. Luna runs $1 and $6. Sol, the flagship, sits at $5 and $30, the same price as GPT-5.5. If cost per token were the only variable, this would be a clean win for OpenAI. It isn't the only variable, and that's the part getting buried under the benchmark chart.

What Actually Shipped on July 9

Rollout is staged, not simultaneous. Plus, Pro, Business, and Enterprise users get Sol through medium and higher reasoning effort settings in ChatGPT. Free-tier users are routed to Terra. Pro and Enterprise accounts can additionally select GPT-5.6 Pro for the hardest tasks. Codex and the API got access on the same day, and GitHub Copilot added all three models within hours of the public launch, with Sol reserved for Pro+, Max, Business, and Enterprise Copilot seats.

The headline feature isn't a benchmark number, it's Ultra mode. Sol can now spin up four subagents in parallel by default, trading extra token spend for speed on long, multi-step jobs. Pair that with a new max-reasoning setting for genuinely hard sequential problems, and you get four operating modes per model, times three models, plus Ultra as a fourth Sol-only layer. That's a lot of knobs for a solo builder to learn, and most of them exist because Sol alone can't reliably out-think Fable 5 on raw intelligence, so OpenAI is compensating with orchestration.

The Benchmark OpenAI Led With, and the One It Didn't

Here's the number from OpenAI's own launch post: Sol hits 53.6 on Agents' Last Exam, a new high, 13.1 points above Fable 5 in adaptive mode. At medium reasoning, Sol still beats Fable 5 by 11.4 points at roughly a quarter of the estimated cost. On the Artificial Analysis Coding Agent Index, Sol scores 80.0, edging Fable 5 by 2.8 points while burning less than half the output tokens and taking less than half the time.

Those are real numbers, and they're specifically about agentic task completion and coding throughput, not general intelligence. That distinction matters more than it sounds. Techzine's independent coverage of the same launch reports Artificial Analysis's general intelligence scores as 59 for Sol, 55 for Terra, and 51 for Luna, landing the family "almost exactly between Fable 5 and Sonnet 5." Read the next line and the picture flips: "according to independent measurements, Fable 5 remains the strongest in general intelligence." OpenAI picked its two best charts. Nobody's lying, but nobody's showing you the whole board either, and that's the gap this article exists to close.

The cybersecurity numbers tell a similar two-sided story. Sol scored 96.7% on OpenAI's internal cyberattack evaluation and matched Anthropic's classified-tier Mythos Preview on ExploitBench at roughly a third of the inference cost, according to reporting confirmed across multiple outlets. That capability jump is exactly why the White House's Office of the National Cyber Director asked OpenAI to hold the June 26 launch to government-vetted partners for 12 days before the July 9 wide release. All three GPT-5.6 models, not just Sol, are now classified at OpenAI's "High" risk level for both cyber and biological or chemical capability. If you're building anything security-adjacent on top of this API, that classification changes your compliance obligations, not just your model choice.

Pricing: Where GPT-5.6 Actually Wins Clean

Model

Input ($/1M tokens)

Output ($/1M tokens)

Best for

GPT-5.6 Sol

$5.00

$30.00

Long-horizon coding, cyber research, max-reasoning tasks

GPT-5.6 Terra

$2.50

$15.00

High-volume production work, ~2x cheaper than GPT-5.5 at comparable output

GPT-5.6 Luna

$1.00

$6.00

Fast, low-cost summarization and routine automation

This table merges OpenAI's own pricing page with VentureBeat's and Techzine's independent coverage of the same launch, since no single source lists all three tiers next to a cost-per-task comparison. On raw dollars, this isn't close. Terra beating GPT-5.5 on performance at half the price is the kind of margin that actually changes a solo creator's monthly API bill, especially if you're running agentic pipelines that chew through output tokens on every retry. Luna at $1/$6 is cheap enough to run as a first-pass filter model in front of a more expensive one, which is exactly the kind of tiered setup worth testing if you're building content or research pipelines on a budget.

New prompt caching rules ship alongside the pricing: explicit cache breakpoints, a 30-minute minimum cache life, and a 90% discount on cached-input reads. Cache writes now cost 1.25x the uncached input rate. If your workflow reuses the same system prompt or long context window across many calls, this caching change is worth more to your actual bill than the headline benchmark gap.

Design Judgment and Computer Use: The Underreported Upgrade

The part of this launch that will matter most to anyone shipping presentations, documents, or spreadsheets isn't the benchmark chart, it's that Sol can now inspect its own rendered output. Earlier agentic models generated code or content and stopped. Sol's stronger computer-use loop lets it look at what actually rendered, catch a misaligned table or a broken chart before handing the file back, and fix it without being asked twice. OpenAI frames this as a "step change in design judgment," and for once the framing matches the mechanism: the model isn't smarter at writing HTML, it's better at checking its own work against what a human would actually see.

That improvement compounds with better template handling in exported artifacts. Presentations, documents, and spreadsheets built through GPT-5.6 are meant to slot into existing enterprise templates and export cleanly into the tools teams already use, rather than needing a manual cleanup pass after every generation. If you've ever asked a model to build a slide deck and gotten back something that technically matched the brief but looked wrong on screen, this is the specific failure mode OpenAI is targeting.

GPT-5.6 vs Claude Fable 5: The Straight Answer

Neither model is a clean win. GPT-5.6 Sol is faster, cheaper, and measurably ahead of Fable 5 on agentic task completion and coding throughput, the two areas OpenAI benchmarked hardest before launch. Claude Fable 5 holds the edge on general intelligence by independent measurement, which matters more for open-ended reasoning, research synthesis, and tasks that don't reduce cleanly to a pass or fail agentic score. If your workload is agent-heavy coding on a budget, Sol's cost-to-performance ratio is hard to argue with. If your workload leans on broad reasoning quality across ambiguous, non-coding tasks, the benchmark data OpenAI didn't lead with still favors Fable 5.

Frequently Asked Questions

Is GPT-5.6 better than Claude Fable 5?

It depends on the task. GPT-5.6 Sol beats Fable 5 on Agents' Last Exam and the Artificial Analysis Coding Agent Index, both agentic and coding-specific benchmarks, at a fraction of the cost. On Artificial Analysis's general intelligence index, independent measurements still favor Fable 5. There is no single "better" answer across both benchmark families.

How much does GPT-5.6 cost compared to Claude Fable 5?

GPT-5.6 Sol costs $5 per million input tokens and $30 per million output tokens, the same as GPT-5.5. Terra runs $2.50/$15 and Luna runs $1/$6. OpenAI's own comparisons put Sol's coding-agent cost at roughly a third less than Fable 5, and Terra and Luna at around one-sixteenth the cost on Agents' Last Exam tasks.

What is Ultra mode in GPT-5.6?

Ultra is Sol's highest-performance setting. It coordinates multiple subagents working in parallel on the same task, trading higher token consumption for faster completion on demanding, long-horizon work.

Is GPT-5.6 available in ChatGPT yet?

Yes, as of July 9, 2026, though the rollout is staged by account tier over roughly 24 hours. Plus, Pro, Business, and Enterprise users get Sol at medium and higher reasoning effort. Free users are routed to Terra. Pro and Enterprise users can also select GPT-5.6 Pro.

Why was GPT-5.6 delayed by 12 days before its public launch?

The White House's Office of the National Cyber Director and Office of Science and Technology Policy asked OpenAI to limit the June 26 launch to roughly 20 government-vetted partners first, citing Sol's cybersecurity capabilities, which scored 96.7% on OpenAI's internal cyberattack evaluation. The gated preview ended July 9 with the full public rollout.

Full technical detail on the launch, including the complete benchmark methodology, is in OpenAI's official announcement.

Keep Reading