GPT-Realtime-2: GPT-5-Class Reasoning Lands in OpenAI's Voice API

GPT-Realtime-2: GPT-5-Class Reasoning Lands in OpenAI's Voice API

5 min readMay 17, 2026

Quick verdict

GPT-Realtime-2 is the first OpenAI voice model where the bottleneck stops being the model and starts being your harness. 96.6% on Big Bench Audio. Context jumps from 32K to 128K. Tool-call transparency, preambles, and adjustable reasoning effort all land in the same release. Audio pricing did not move: $1.15/hr in, $4.61/hr out.

What actually shipped

Three models in the Realtime API today: GPT-Realtime-2 for voice-to-voice agents, GPT-Realtime-Translate for live speech translation across 70+ input and 13 output languages, and GPT-Realtime-Whisper for low-latency streaming transcription. The headline model brings GPT-5-class reasoning to a native speech-to-speech architecture.

  • Context window expanded from 32K to 128K tokens, with 32K max output
  • Reasoning effort selectable across minimal, low, medium, high, and xhigh. Default is low
  • Time-to-first-audio of 1.12s at minimal reasoning and 2.33s at high
  • Preambles like "let me check that" and audible tool-call narration like "checking your calendar"
  • Parallel tool calls and recovery phrases when something fails
  • Better tone control, domain vocabulary retention, and quieter behavior when the user is clearly talking to someone else in the room
  • Inputs include text, audio, and image

On benchmarks, Artificial Analysis measured 96.6% on Big Bench Audio speech-to-speech reasoning and 96.1% on Conversational Dynamics. Scale AI placed it first on the Audio MultiChallenge S2S leaderboard with instruction retention climbing from 36.7% to 70.8% APR over Realtime-1.5. The customer numbers back it: Glean reported a 42.9% relative helpfulness lift in internal evals, Genspark's Call for Me agent reported a 26% increase in effective conversation rate.

Why it matters

Realtime-1.5 was a 4o-class bump, +5% on Big Bench Audio. Realtime-2 lands +15.2% and closes the reasoning gap with text models. That's what was missing for voice agents that have to handle tool calls and recover from interruptions inside a single session.

The 128K context resets the unit economics. An agent can hold a full support transcript and a tool schema without losing the thread. Pricing held flat, so for anyone already on the Realtime API, the upgrade is free compute.

The one caveat: ChatGPT Voice itself has not been upgraded. If you were expecting the consumer app to feel different today, it doesn't. The win is API-only for now.

Video: GPT-Realtime-2 demo

OpenAI's Build Hour walkthrough covers the new model's reasoning, tool calls, and translation flow end-to-end.

FAQ

Is GPT-Realtime-2 available in ChatGPT Voice yet?

No. The models are live in the Realtime API only. Sam Altman said ChatGPT voice improvements are still cooking.

How much does it cost?

$1.15 per hour of audio input and $4.61 per hour of audio output, unchanged from the previous Realtime model. The 4x context expansion and reasoning upgrade are essentially free for anyone already on the API.

How does it compare to Whisper for transcription?

GPT-Realtime-Whisper is the streaming variant built for low-latency captions and continuous understanding inside the Realtime API, rather than batch file transcription. Plain Whisper still wins for offline processing.

What languages does Realtime-Translate cover?

Streaming speech translation from 70+ input languages into 13 output languages.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles