Learning Objectives
- Understand how GPT-5.5 differs from GPT-5.4 and where it changes the practical defaults
- Compare the three GPT-5.5 variants (standard, Thinking, Pro) and identify the right one for different use cases
- Evaluate GPT-5.5 against the frontier models it launched against (Claude Opus 4.7, Gemini 3.1 Pro)
What Is GPT-5.5?
GPT-5.5 was OpenAI's flagship general-purpose model from its release on April 23, 2026 — just six weeks after GPT-5.4 — until GPT-5.6 replaced it on July 9, 2026. OpenAI described it as "a new class of intelligence built specifically for real work and for powering agents," and it replaced GPT-5.4 as the default model in ChatGPT for Plus, Pro, Business, and Enterprise users.
📝Note
Now the recent prior generation. OpenAI shipped GPT-5.6 (Sol, Terra, Luna) to general availability on July 9, 2026, and it is now the flagship — see the dedicated GPT-5.6 page. GPT-5.5 remains widely available and is still an excellent agentic model; the details below reflect its April launch.
The headline shift from 5.4: GPT-5.5 is engineered for agentic workflows — multi-step tasks where the model plans, uses tools, executes commands, checks its own output, and recovers from mistakes with fewer human iterations. It also uses significantly fewer tokens than GPT-5.4 on the same Codex tasks at matched latency, which lowers the cost of long agent runs.
✅Tip
Access GPT-5.5: Available in ChatGPT for Plus, Pro, Business, and Enterprise users, and through the OpenAI API, which completed its staged rollout with separate safeguards for autonomous use. GPT-5.5, GPT-5.4, and Codex are also generally available on Amazon Bedrock as of June 2026, having graduated from the earlier limited preview — the first time OpenAI frontier models have shipped through a cloud other than Microsoft Azure since 2019, after the Microsoft-OpenAI agreement amendment ended exclusive cloud rights through 2032. GPT-5.5 runs in AWS's US East region and GPT-5.4 in US East and US West, callable through the Responses API at pricing that matches OpenAI's first-party rates.
Key Capabilities
Agentic Coding and Computer Use
GPT-5.5's biggest jumps over GPT-5.4 are on the benchmarks that measure real agent work:
- Terminal-Bench 2.0: 82.7% (GPT-5.4: 75.1%) — multi-tool command-line workflows requiring planning and error recovery; the metric most coding-agent builders are paying attention to
- OSWorld-Verified: 78.7% (GPT-5.4: 75.0%) — desktop UI automation across mixed software tasks
- BrowseComp: 90.1% (GPT-5.5 Pro) — web research agent capability for locating hard-to-find information
The token-efficiency gain compounds with capability: agents that used to burn through context on retries finish the same task in fewer steps.
Knowledge Work
- GDPval: 84.9% (GPT-5.4: 83.0%) — economically valuable knowledge-work tasks measured against expert humans
- Stronger handling of unstructured prompts: GPT-5.5 plans workflows from a one-line goal where GPT-5.4 needed scaffolding
Scientific and Mathematical Reasoning
- FrontierMath Tier 4: 35.4% (Claude Opus 4.7: 22.9%, Gemini 3.1 Pro: 16.7%) — the hardest open math benchmark, where GPT-5.5 led frontier models by a wide margin at launch
- Meaningful gains on technical research workflows; OpenAI cites drug-discovery scenarios as an early use case
Context Window
GPT-5.5 carries forward a roughly 1 million token context window (922K effective) for the standard API. The Codex variant ships with a 400K context tuned for long-running coding sessions.
GPT-5.5 Variants
| Variant | Best For | Availability |
|---|---|---|
| GPT-5.5 Instant | Default ChatGPT model across every tier as of May 5, 2026; reduced hallucinations on law / medicine / finance; AIME 2025 81.2 vs prior 65.4 | Free; Plus; Pro; Business; Enterprise; API as `chat-latest` |
| GPT-5.5 Thinking | Complex reasoning requiring extended chains of thought | Plus; Pro; Business; Enterprise |
| GPT-5.5 Pro | Maximum capability; highest reasoning effort per request; 90.1% BrowseComp | Pro; Business; Enterprise |
Choosing between variants:
- Default to GPT-5.5 Instant for most tasks — same low-latency profile users expect from chat, now with materially fewer hallucinations on sensitive professional domains
- GPT-5.5 Thinking when the task requires extended reasoning (complex math, multi-step logic, hard coding tickets)
- GPT-5.5 Pro when accuracy on the hardest research, browsing, or computer-use tasks justifies the higher compute cost
May 5, 2026 — Instant becomes default across every tier
OpenAI swapped ChatGPT's default model from GPT-5.3 Instant to GPT-5.5 Instant for free, Plus, Pro, Business, and Enterprise users on May 5, 2026, while keeping the same low-latency profile. The company says the new model substantially reduces hallucinations in sensitive professional domains (law, medicine, finance) and posts an AIME 2025 math score of 81.2 versus the prior 65.4, alongside improvements in multimodal reasoning. GPT-5.3 will remain available to paid API users for three more months before the older default is fully retired. The Instant variant is exposed to API users immediately as chat-latest; named gpt-5.5 and gpt-5.5-pro endpoints continue their separate rollout.
OpenAI did not ship a mini or nano variant alongside the GPT-5.5 launch. Older 5.x models (5.1, 5.2, 5.3, plus the existing GPT-5.4 mini) remain on the menu for cost-sensitive serving.
GPT-5.5-Cyber: Defensive Security Variant
On June 22, 2026, OpenAI fully released GPT-5.5-Cyber, a variant tuned specifically for defensive cybersecurity work as part of its expanded Daybreak security program. OpenAI calls it its strongest model yet for finding, validating, and patching software vulnerabilities, pointing to its ability to sustain deeper analysis across large codebases. It set a new state of the art on the CyberGym benchmark at 85.6 percent, up from 81.8 percent for the standard GPT-5.5.
Alongside the model, OpenAI launched Patch the Planet, an initiative that funds expert security researchers — working with the firms Trail of Bits and HackerOne — to fix flaws in widely used open-source projects and collaborate directly with their maintainers; more than 30 projects have signed on. OpenAI also opened GPT-5.5 to roughly 30 cybersecurity vendors to embed in their own defensive products. The company's framing is that its models now surface vulnerabilities faster than defenders can fix them, so the bottleneck has shifted from finding bugs to patching them.
Pricing
ChatGPT Subscription Tiers
- GPT-5.5 Instant default (May 5, 2026)
- Usage-limited
- Thinking and Pro variants not included
- GPT-5.5 Instant default with higher limits
- Thinking and Pro variants not included
- GPT-5.5 Instant + GPT-5.5 Thinking
- Generous limits
- GPT-5.5 Pro
- Unlimited GPT-5.5
- Maximum compute and memory
- GPT-5.5 + Thinking + Pro
- Team workspace and admin controls
- Full GPT-5.5 + Pro access
- Data privacy
- SSO and compliance
API Pricing
Third-party reporting at launch (Apr 23, 2026) cites the following per-million-token pricing for the upcoming gpt-5.5 and gpt-5.5-pro API endpoints:
- Standard agentic model
- Roughly twice the per-token cost of gpt-5.4
- Pending official OpenAI publication
- Pro reasoning effort
- Highest accuracy on hardest tasks
- Pending official OpenAI publication
- Higher latency tolerance
- Same model quality
- Pending official OpenAI publication
The named endpoints have since gone live; check platform.openai.com/pricing for the current official rate card, which supersedes the launch-day reporting above.
GPT-5.5 vs. the Frontier Models of Its Generation
| Benchmark | GPT-5.5 | GPT-5.4 | Claude Opus 4.7 | Gemini 3.1 Pro |
|---|---|---|---|---|
| Terminal-Bench 2.0 | 82.7% | 75.1% | 69.4% | 68.5% |
| GDPval (knowledge work) | 84.9% | 83.0% | 80.3% | 67.3% |
| OSWorld-Verified | 78.7% | 75.0% | 78.0% | — |
| SWE-bench Pro | 58.6% | 57.7% | 64.3% | 54.2% |
| FrontierMath Tier 4 | 35.4% | — | 22.9% | 16.7% |
Against that field GPT-5.5 led on the agentic and mathematical benchmarks (Terminal-Bench, GDPval, OSWorld, FrontierMath), while Claude Opus 4.7 led on SWE-bench Pro at 64.3%. That last gap has only widened: Anthropic's Claude Fable 5 now leads SWE-bench Pro at 80.3% against GPT-5.5's 58.6%, and OpenAI has not published a SWE-bench Pro score for GPT-5.6. Pick a Claude model when you need one that can grind through complex GitHub issues end-to-end.
Strategic Context: The "Super App"
GPT-5.5 is the model OpenAI's leadership has framed as a step toward a "super app" — a single product that bundles ChatGPT, Codex, and the AI browser into one assistant that can carry a task across tools without handing it off. The fact that 5.5 ships only six weeks after 5.4 underscores how aggressively OpenAI is iterating to defend share against Claude and Gemini in the enterprise market.
Strengths
- Strong agentic coding — its Terminal-Bench 2.0 score of 82.7% led the field at launch and remains competitive
- Strong knowledge-work model — GDPval 84.9%, ahead of Claude Opus 4.7 and Gemini 3.1 Pro in its generation
- Token-efficient — uses fewer tokens than GPT-5.4 for the same Codex tasks at matched latency, lowering long-agent-run cost
- Scientific reasoning leap — FrontierMath Tier 4 35.4% nearly doubled Claude Opus 4.7 on the hardest open math benchmark
- Largest developer ecosystem — same OpenAI API, SDKs, tutorials, and community as the rest of the GPT-5 family
Limitations & Considerations
- No longer the flagship — GPT-5.6 replaced it on July 9, 2026; start new work on GPT-5.6 unless you have a specific reason to pin to this version
- Higher API price than GPT-5.4 — reported at $5 in and $30 out per million tokens, roughly twice the cost of
gpt-5.4; recheck math before migrating high-volume jobs - Free and Go tiers do not include GPT-5.5 — those tiers stay on older 5.x mini models
- Claude leads SWE-bench Pro by a wide margin — Claude Fable 5 posts 80.3% against GPT-5.5's 58.6%, so for the hardest end-to-end software-engineering tasks a Claude model is the safer default
- Closed model — API-only access; no weights available for self-hosting
- Rapid release cadence — OpenAI shipped 5.4 and 5.5 within six weeks; assume the model behind your prompts can change again on a similar timeline
Succeeded by GPT-5.6 (Sol, Terra, Luna)
On July 9, 2026, OpenAI made GPT-5.6 generally available across ChatGPT, Codex, and the API — it is now the flagship. It introduces a new naming scheme: the number marks the generation, while Sol, Terra, and Luna name durable capability tiers that can advance on their own cadence. GPT-5.6 also adds a max reasoning effort setting and an "ultra mode" that coordinates subagents to work cooperatively rather than as a single agent — the top "Sol Ultra" configuration is coming to OpenAI's Codex coding agent.
| Tier | Role | API price (per million tokens) |
|---|---|---|
| Sol | Flagship, deepest reasoning | $5 in / $30 out |
| Terra | Balanced lower-cost option | $2.50 in / $15 out |
| Luna | Fastest, most cost-efficient | $1 in / $6 out |
For the full breakdown — including the honest benchmark caveats (all scores vendor-reported; no published SWE-bench Pro number) — see the dedicated GPT-5.6 page. GPT-5.5 remains widely available as the recent prior generation.
Lineage
- GPT-5.6 (generally available July 9, 2026) — current flagship family; Sol, Terra, and Luna tiers plus an ultra mode. See the GPT-5.6 page
- GPT-5.5 (April 23, 2026) — recent prior flagship; the page above
- GPT-5.4 (March 5, 2026) — prior flagship; introduced 1 million token context, native computer-use, and the GPT-5.4 Pro / Thinking / mini / nano variant lineup. Still available via the
gpt-5.4API endpoint and as the underlying Free / Go tier model - GPT-5.3-Codex — dedicated coding model that powers the OpenAI Codex platform
- GPT-5.3-Codex-Spark — real-time coding variant on Cerebras hardware (1,000+ tokens/sec)
- GPT-5.5-Cyber (June 22, 2026) — defensive security variant for finding and patching vulnerabilities; CyberGym 85.6%; anchors the Daybreak program and Patch the Planet
- GPT-OSS — OpenAI's open-weight model under Apache 2.0
- GPT-5.1 was deprecated on March 11, 2026; GPT-5.2 and GPT-5.3 remain on the menu for cost-sensitive serving
Key Takeaways
- GPT-5.5 (April 23, 2026) was OpenAI's flagship until GPT-5.6 replaced it on July 9, 2026 — built specifically for agentic work, with launch-record scores on Terminal-Bench 2.0 (82.7%), GDPval (84.9%), OSWorld (78.7%), and FrontierMath Tier 4 (35.4%). It remains widely available as the recent prior generation
- Three variants ship: Instant (May 5, 2026 default ChatGPT model across every tier; reduced hallucinations on law / medicine / finance; AIME 2025 81.2 vs prior 65.4), Thinking, and Pro — no mini or nano this time
- The Instant variant is live in the API as
chat-latest, and the namedgpt-5.5andgpt-5.5-proendpoints have since completed their separate rollout at a reported $5 in and $30 out per million tokens; GPT-5.3 Instant stayed available to paid API users for three months after the swap - Now generally available on Amazon Bedrock (June 2026, after an earlier limited preview) — the first non-Azure cloud distribution of OpenAI frontier models since 2019, after the Microsoft-OpenAI agreement amendment ended exclusive cloud rights through 2032
- Anthropic leads SWE-bench Pro — Claude Fable 5 at 80.3% against GPT-5.5's 58.6%, and OpenAI has published no SWE-bench Pro score for GPT-5.6 — so pick a Claude model for the hardest end-to-end software-engineering tasks