← Back to all posts
News

OpenAI's GPT-5.6 Sol: New Records, an 'Ultra' Subagent Mode, and a Third of the Tokens

June 26, 2026 · News
OpenAI's GPT-5.6 Sol: New Records, an 'Ultra' Subagent Mode, and a Third of the Tokens

TL;DR

OpenAI previewed GPT-5.6 as a three-model family: Sol (the flagship), Terra (balanced, cheaper), and Luna (fastest and cheapest). The headline is not just higher scores, it is two new gears: a max reasoning effort that lets Sol think longer, and an ultra mode that goes beyond a single agent by spinning up subagents to attack complex work in parallel. Sol in ultra mode posts a record 91.9% on Terminal-Bench 2.1, and across biology and cybersecurity it matches or beats rivals while burning a fraction of the tokens. The asterisk: at the US government's request, GPT-5.6 is launching as a limited preview to roughly 20 approved partners, so most builders cannot touch it yet.

Terminal-Bench 2.1: command-line agent tasks (higher is better) Sol (ultra)91.9 Sol (max)88.8 Mythos 588.0 Fable 584.3 GPT-5.583.4
Even base Sol clears the Anthropic frontier models; ultra mode pulls further ahead. Scores per OpenAI.

The lineup: Sol, Terra, Luna

OpenAI is doing the now-familiar three-tier split, so you pick the model that fits the job instead of paying flagship rates to fix a typo. Pricing is per 1M tokens.

ModelRoleInput / Output
Solflagship, hardest work$5 / $30
Terrabalanced, everyday$2.50 / $15
Lunafast and cheap, high volume$1 / $6

What "ultra" mode actually does

This is the genuinely new idea. A normal model answers as one agent: it plans, calls tools, and works the problem itself, step by step. Ultra mode lets Sol act less like a single worker and more like a manager. It spins up subagents, hands each a slice of a complex task, lets them work in parallel, and then stitches the results back together.

Picture a head chef during a dinner rush. A line cook works one dish at a time, start to finish. The head chef does not cook every plate, they direct a brigade: one cook on sauces, one on the grill, one plating, all going at once, so a complicated three-course order leaves the kitchen far faster than any single pair of hands could manage. Ultra mode is Sol promoting itself from line cook to head chef. That is how it clears 50.9% on Agent's Last Exam in code mode, the first model to get past the halfway mark on that benchmark.

single agent vs. ultra mode NORMAL one agent, step by step ULTRA Sol subagent A subagent B subagent C merged
Ultra mode turns one model into an orchestrator running parallel copies of itself.

The efficiency story is the real flex

Records are nice, but the number builders should notice is token cost. OpenAI says Sol hits its stronger biology and genomics results while consuming fewer tokens than GPT-5.5 and GPT-5.4, and on ExploitBench it stays competitive with Anthropic's Mythos Preview using only about one third of the output tokens. Output tokens are where the bill lives, so a model that matches a rival on a third of them is quietly a pricing story as much as a capability one. Strong on coding, science, and cybersecurity, with vulnerability research and exploitation called out specifically, and the efficiency to run those long-horizon tasks without the token meter spinning out.

The catch: you probably cannot use it

Here is where the excitement meets reality. GPT-5.6 is not a normal launch. At the US government's request, OpenAI is starting with a limited preview for a small group of trusted partners (reported as around 20 companies) whose participation has been shared with the government, citing the model's advanced autonomous cybersecurity capability. General availability for Sol, Terra, and Luna is promised "in the coming weeks." We unpacked the policy mechanics, the executive order, and the customer-by-customer approval in a separate post; for this one the practical point is simpler: the model is real and strong, but unless you are on the list, you are reading benchmarks you cannot yet reproduce.

The honest caveats

Every number here is OpenAI's own, on OpenAI's chosen benchmarks, published the same day as the model, with no independent replication yet. Terminal-Bench and Agent's Last Exam are real and respected, but a vendor leaderboard on launch day is a marketing artifact until outside labs run it. The token-efficiency claims are the most interesting and the least verifiable, so treat them as a promise to test, not a settled fact. And ultra mode's subagent orchestration sounds powerful, but parallel subagents can multiply cost and failure modes as easily as they multiply throughput. When you can finally get in, benchmark it on your own workload before you believe the chart.

Key Takeaways

  • GPT-5.6 ships in three sizes: Sol (flagship, $5/$30 per 1M tokens), Terra ($2.50/$15), and Luna ($1/$6).
  • Two new gears: a max reasoning effort for deeper single-agent thinking, and an ultra mode that orchestrates subagents in parallel.
  • Sol sets a record 91.9% on Terminal-Bench 2.1 in ultra mode, and is the first model past the halfway mark (50.9%) on Agent's Last Exam in code mode.
  • The efficiency is the sleeper story: Sol matches Anthropic's Mythos Preview on ExploitBench using about a third of the output tokens.
  • It is a limited preview to roughly 20 government-approved partners, with general availability promised in the coming weeks. Verify the numbers yourself once you can.

Sources: OpenAI: Previewing GPT-5.6 Sol, OpenAI GPT-5.6 Preview System Card

AIOpenAIGPT-5.6Solbenchmarksagentscodingcybersecurity
CONSOLE
$