
GPT-5.6 Is Now Live for Everyone, and OpenAI Turned the Launch Into a Superapp
Quick verdict
The gated preview is over. GPT-5.6 shipped to everyone across ChatGPT, Codex, and the API, and OpenAI used the moment to launch a stack of products around it: ChatGPT Work, a single desktop app that folds Codex into ChatGPT, and Sites. On the model itself, the flagship Sol lands second to Claude Fable 5 on broad intelligence but costs about a third as much per task, which is the whole pitch. The catch is a safety one. The UK's AI Security Institute says it found universal jailbreaks in every round of testing.
What actually shipped
OpenAI released three models. Sol is the flagship with the highest reasoning ceiling for long coding and agent work, Terra is the balanced middle, and Luna is the fast, cheap tier for high volume. All three roll out across ChatGPT, Codex, and the API. Headline API pricing matches GPT-5.5, so you get more capability at the same sticker.
| Model | Tier | API price (per 1M in / out) | AA Intelligence Index | Cost per task |
|---|---|---|---|---|
| Sol | Flagship | $5 / $30 | 59 | $1.04 |
| Terra | Balanced | $2.5 / $15 | 55 | $0.55 |
| Luna | High-volume | $1 / $6 | 51 | $0.21 |
One pricing wrinkle worth flagging: OpenAI added cache-write pricing for the first time, charging cache writes at 1.25 times the input token price. Cache reads keep the familiar 90 percent discount. If you build agents that write a lot of context to cache, this changes the math.
The product launch was bigger than the model. OpenAI shipped:
- ChatGPT Work, an agent powered by Codex and GPT-5.6 that operates across your apps and files and can stay on a project for hours.
- A merged desktop app that pulls Codex and ChatGPT into one surface, with browser integration, Chrome extension support, faster Computer Use, and shared context between Work and Codex.
- Sites in beta, which turns outputs into shareable web pages.
- On the API, Programmatic Tool Calling in the Responses API and a multi-agent mode in beta. The desktop stack now handles authenticated sites, multi-tab sessions, and file downloads.
How it benchmarks
Independent numbers landed fast, and they tell a consistent story: near the top on capability, clearly ahead on price. Artificial Analysis put Sol at 59 on its Intelligence Index, a close second to Claude Fable 5, while costing roughly a third as much per task. On the Coding Agent Index, Sol scored 80 and led the board, with Terra at 77 and Luna at 75. Sol running in Codex topped all three of the index's coding evaluations: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
A few standout results:
- Arc Prize said Sol set a new record on ARC-AGI-3 at 7.8 percent and became the first verified frontier model to beat an ARC-AGI-3 game, up from Opus 4.8's 1.5 percent. One caveat: a widely shared thread argued that under the official $10k budget cap Sol would have scored zero, since it was reportedly allowed $25k.
- On ARC-AGI-2, Sol hit 92.5 percent at one order of magnitude lower cost than GPT-5.5 Pro.
- Vals ranked GPT-5.6 second on its main and multimodal indexes, with Sol taking first on CyberBench, Excel Modeling, Legal Research, ProofBench, SWE-bench, and Terminal-Bench 2.1.
- Cursor reported Sol at 67.2 percent on CursorBench.
Not everything moved forward. Artificial Analysis called the gain on its AA-Omniscience test minor over GPT-5.5, with a small rise in hallucination rate. Some testers flagged that Sol may be weaker at math. And Artificial Analysis argued Terra is not on the cost-versus-intelligence frontier at all, because there is usually a Sol or Luna operating point that matches it for the same money.
Why it matters
OpenAI is not claiming to be the smartest model anymore. The framing this week was dollars per task, not benchmark supremacy. Sam Altman said the company heard enterprise complaints about AI costs and that Sol is a big step for dollars per task. That is a real shift. Once the top models sit within a few points of each other on coding and agent work, the deciding factor becomes which one finishes the job with the least token spend and wall-clock time.
This puts routing squarely on the table. With Sol, Terra, Luna, Claude Fable 5, Grok 4.5, and Meta's new Muse Spark 1.1 all sitting on different points of the cost curve, no single model is the right default for every request. That is exactly the problem that cost-based model routing and routing your coding agent are built to solve, and it is why paying for one flagship subscription while ignoring the cheaper tiers leaves money on the table.
There is also a product story. Bundling ChatGPT Work, a merged desktop app, and Sites with the model turns a launch into an operating environment for AI work. Not everyone loves the consolidation. One well-known developer called folding Codex into ChatGPT Desktop a generational fumble, worried the focused coding experience gets diluted. Inside ChatGPT subscriptions, Sol also burns twice the credits of GPT-5.5, so the cheap-per-token story does not fully carry over to the flat-rate plans most people actually use.
The safety problem nobody can wave away
OpenAI called GPT-5.6 its most capable model yet on cyber and bio tasks, and said some API calls may be paused mid-stream for review in dual-use areas. Then the UK's AI Security Institute published its findings. In every round of testing, it reported universal jailbreaks that let the model complete long agentic tasks in areas including vulnerability discovery and exploit development. One safety researcher called it the highest-stakes safety issue of any model release yet. Part of what lifts Sol's scores on some evals is a lower refusal rate than competitors, which is useful for real work and also raises the misuse ceiling.
Video: GPT-5.6 and OpenAI's superapp push
A walkthrough of the GPT-5.6 launch and the product bundle around it, from the model tiers to ChatGPT Work and Sites.
FAQ
Is GPT-5.6 available to everyone now?
Yes. Unlike the earlier gated preview covered in our restricted rollout piece, the family is now rolling out across ChatGPT, Codex, and the API, expanding over the first 24 hours.
Which GPT-5.6 model should I use?
Sol for the hardest long-horizon coding and agent runs, Luna for cheap high-volume work. Terra sits in the middle but independent analysis suggests a Sol or Luna setting often matches it for the same cost, so test before defaulting to it.
How does it compare to Claude Fable 5?
Fable 5 still leads on broad intelligence, but Sol runs at roughly a third of the cost per task and tops several coding and agent evals. See our Claude Fable 5 launch coverage for the other side.
Did GPT-5.6 really train another model on its own?
OpenAI's claim that Sol autonomously post-trained Luna was the most-discussed line of the launch, but researchers pushed back that the real scope was narrower, closer to Sol editing a config and launching a run in a controlled setup than end-to-end training.
Sources
- @OpenAI - GPT-5.6 Sol, Terra, and Luna rollout across ChatGPT, Codex, and the API
- @sama - Sol as a step forward for dollars per task, alongside Terra and Luna
- @ArtificialAnlys - Intelligence Index, coding index, and per-task cost breakdown
- @OpenAI - ChatGPT Work, the new Codex + GPT-5.6 agent
- @OpenAIDevs - merged Codex and ChatGPT desktop app
- @arcprize - Sol sets a new ARC-AGI-3 record at 7.8%
- @ValsAI - GPT-5.6 rankings on the Vals Index and Multimodal Index
- @cursor_ai - GPT-5.6 in Cursor, Sol at 67.2% on CursorBench
- @alxndrdavies - UK AI Security Institute universal jailbreak findings
- @theo - calling the Codex-into-ChatGPT Desktop merge a generational fumble
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix