The OpenAI Realtime API lets an application hold a live, two-way voice conversation with a speech-to-speech model. Instead of chaining a transcriber, a text model and a speech synthesiser, you stream microphone audio in and receive synthesised audio out over one long-lived connection, with the conversation state, turn detection and tool calls managed by the session. That removes two hand-offs from the latency budget, and it changes how you build: everything is an event on a socket rather than a request and a response.
This article explains the moving parts from first principles: the three transports, ephemeral credentials, the session and event model, turn detection, barge-in with truncation, function calling, the audio arithmetic, and how a hosted speech-to-speech session compares with running a voice pipeline on your own GPUs. Names below are checked against OpenAI's current documentation and its Python SDK type definitions as of October 2026; the API is evolving, so pin your SDK and recheck event names when you upgrade.
Why speech-to-speech changes the latency budget
A cascaded voice agent waits for the user to stop talking, finalises a transcript, sends text to a language model, waits for enough tokens to synthesise a phrase, then streams audio. The user-perceived delay is the sum of endpointing time, transcription finalisation, model time to first token, synthesis time to first audio and network hops between three services. Each stage can stream, but each boundary adds buffering, and tone, hesitation and emphasis are lost the moment speech becomes text.
A speech-to-speech model consumes audio directly and emits audio directly. Endpointing still exists (the model has to decide the user has finished) and so does time to first output, but the transcription and synthesis hand-offs disappear from the critical path. The price is control: you cannot swap the transcriber, insert a text rewrite before speech, or pick a different synthesis engine per language. If you need those, a cascade is still the right design.
Transports and credentials
| Transport | Best for | Audio path | Events |
|---|---|---|---|
| WebRTC | Browsers and mobile apps | Media tracks; the browser encodes and handles jitter and echo | JSON over a data channel named oai-events |
| WebSocket | Server-to-server, custom audio pipelines | Base64 audio inside JSON events | Same JSON events on the socket |
| SIP | Telephony via a SIP trunk | Phone-call media | Your server accepts, rejects or refers calls and attaches a session |
Never ship your standard API key to a client. For WebRTC, your backend calls POST /v1/realtime/client_secrets with the session configuration and returns the short-lived value to the device, which uses it to post its SDP offer to /v1/realtime/calls. The expires_after setting accepts 10 to 7,200 seconds and defaults to 600. Keep it short: the secret only has to last until the device connects.
# backend: mint a short-lived client secret (Python, requests)
import os, requests
def mint_client_secret(user_id: str) -> dict:
r = requests.post(
"https://api.openai.com/v1/realtime/client_secrets",
headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
json={
"expires_after": {"anchor": "created_at", "seconds": 60},
"session": {
"type": "realtime",
"model": "gpt-realtime",
"instructions": "You are a concise support agent. Never read out card numbers.",
"audio": {"output": {"voice": "marin"}},
},
},
timeout=10,
)
r.raise_for_status()
body = r.json()
return {"key": body["value"], "expires_at": body["expires_at"]}Configure the session on the server when minting, not on the client. Anything the device sends in session.update is under the user's control, so instructions and tool lists that matter for safety belong in the server-side request.
Sessions, items and events
A session holds configuration and a conversation: an ordered list of items (user messages, assistant messages, function calls and function call outputs). A response is one model turn that appends items. Clients send events to change configuration, add input and ask for responses; the server streams events describing what happened. The ones you will handle in almost every application are these.
| Direction | Event | Meaning |
|---|---|---|
| client | session.update | Change instructions, tools, voice, audio formats, turn detection |
| client | input_audio_buffer.append | Add base64 audio to the input buffer |
| client | input_audio_buffer.commit | Close the user turn manually (when turn detection is off) |
| client | conversation.item.create | Insert a text message or a function call output |
| client | response.create / response.cancel | Start or stop a model response |
| client | conversation.item.truncate | Cut an assistant audio item to what was actually played |
| server | input_audio_buffer.speech_started / speech_stopped | Turn detection saw speech begin or end |
| server | response.output_audio.delta | A chunk of base64 output audio |
| server | response.output_audio_transcript.delta | Text of what the model is saying |
| server | response.function_call_arguments.done | A tool call is ready: name, call_id, arguments |
| server | response.done | The response finished; includes status and token usage |
| server | error | A client event was rejected; the session usually continues |
Audio formats are audio/pcm (16-bit little-endian at 24 kHz only), audio/pcmu and audio/pcma (G.711, the formats telephony already uses). Set input and output formats explicitly; a mismatch between what you send and what you declare does not raise an error, it produces noise or chipmunk audio.
A server-side client in Python
The loop below is a server-side WebSocket client. It configures the session, streams microphone frames, plays audio deltas, handles barge-in and runs one tool. Audio capture and playback are abstracted behind mic and player so the protocol logic stays visible; player.stop() flushes queued audio and returns how many milliseconds of the current item actually reached the speaker.
import asyncio, base64, json, os
import websockets # websockets >= 14 uses additional_headers
URL = "wss://api.openai.com/v1/realtime?model=gpt-realtime"
TOOLS = [{
"type": "function", "name": "lookup_ticket",
"description": "Fetch status for a support ticket id.",
"parameters": {"type": "object", "properties": {"ticket_id": {"type": "string"}},
"required": ["ticket_id"]},
}]
async def run(mic, player, lookup_ticket):
hdrs = {"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"}
async with websockets.connect(URL, additional_headers=hdrs) as ws:
await ws.send(json.dumps({"type": "session.update", "session": {
"type": "realtime",
"instructions": "Answer briefly. Use lookup_ticket for ticket questions.",
"tools": TOOLS,
"audio": {
"input": {"format": {"type": "audio/pcm", "rate": 24000},
"turn_detection": {"type": "server_vad", "threshold": 0.5,
"prefix_padding_ms": 300,
"silence_duration_ms": 500}},
"output": {"format": {"type": "audio/pcm", "rate": 24000}, "voice": "marin"},
}}}))
async def pump_mic():
async for frame in mic.frames(ms=20): # 960 bytes per 20 ms frame
await ws.send(json.dumps({"type": "input_audio_buffer.append",
"audio": base64.b64encode(frame).decode()}))
asyncio.create_task(pump_mic())
speaking_item, generating = None, False
async for raw in ws:
ev = json.loads(raw)
t = ev["type"]
if t == "response.created":
generating = True
elif t == "response.output_audio.delta":
speaking_item = ev["item_id"]
player.enqueue(base64.b64decode(ev["delta"]))
elif t == "input_audio_buffer.speech_started" and speaking_item:
played = player.stop() # flush local audio now
if generating:
await ws.send(json.dumps({"type": "response.cancel"}))
await ws.send(json.dumps({"type": "conversation.item.truncate",
"item_id": speaking_item, "content_index": 0, "audio_end_ms": played}))
speaking_item = None
elif t == "response.done":
generating = False
calls = [i for i in ev["response"].get("output") or []
if i["type"] == "function_call"]
for call in calls: # answer every call first
result = await lookup_ticket(**json.loads(call["arguments"]))
await ws.send(json.dumps({"type": "conversation.item.create", "item": {
"type": "function_call_output", "call_id": call["call_id"],
"output": json.dumps(result)}}))
if calls: # then one new response
await ws.send(json.dumps({"type": "response.create"}))
elif t == "error":
print("realtime error:", ev.get("error"))With server VAD and its default create_response behaviour, the server starts a response when it detects the end of speech, so the client never sends response.create for ordinary turns. It does send one after function call outputs, because a tool result is not a user turn. The code waits for response.done before answering calls: response.function_call_arguments.done arrives while the response is still active, and a response may contain several calls, which should all be answered before a single response.create. When interrupt_response is enabled the server also cancels an in-progress response on new speech, so the explicit cancel can race it; an error saying no response is active is harmless. If a reply arrives as a function call while the tool is slow, have the model say a short holding phrase first, or the user hears silence and starts talking over it.
Turn detection
Turn detection decides when the user has finished. server_vad is energy-based: speech above threshold (default 0.5) starts a turn, prefix_padding_ms (default 300) of earlier audio is kept so the first syllable is not clipped, and silence_duration_ms (default 500) of quiet ends it. Shorter silence feels snappier but cuts off people who pause mid-sentence; raise the threshold in noisy rooms. semantic_vad instead uses a model to judge whether the utterance sounds finished, with an eagerness of low, medium, high or auto. It tolerates thinking pauses better at the cost of occasionally waiting longer. For push-to-talk, set turn_detection to null, then send input_audio_buffer.commit and response.create when the button is released.
Barge-in and truncation, worked through
Barge-in is the hardest part to get right and the easiest to get subtly wrong. The server generates audio faster than real time, so when the user interrupts, the conversation history already contains the full assistant answer even though the user heard only part of it. If you do nothing, the model believes it said things the user never heard and will refer back to them. conversation.item.truncate fixes the record: it discards the unplayed audio and its transcript after audio_end_ms.
Worked example. The assistant starts a 6-second answer; your player has written 77,760 bytes of 24 kHz 16-bit mono PCM to the speaker when speech_started arrives. That format runs at 24,000 samples times 2 bytes = 48,000 bytes per second, so 77,760 bytes is 1.62 seconds: send audio_end_ms = 1620. Use the bytes that left the output device, not the bytes received; a 200 ms jitter buffer otherwise makes the model think the user heard a fifth of a second more than they did. Over WebRTC the server manages the output buffer, and output_audio_buffer.clear stops playback that has not reached the device yet.
Audio arithmetic and session limits
The arithmetic matters for capacity and for bugs. PCM16 at 24 kHz is 48,000 bytes per second, 384 kbit/s before encoding; base64 inflates it by a third to 64,000 bytes per second each way on a WebSocket, plus JSON framing. A 20 ms frame is 960 bytes. G.711 is 8 kHz, one byte per sample, 64 kbit/s, which is why telephony integrations should declare audio/pcmu or audio/pcma and skip resampling entirely. WebRTC uses compressed media, which is one reason it is the better choice on mobile networks.
Sessions are long-lived: the documentation currently gives a 60-minute maximum, and every turn's audio accumulates as context. Long calls therefore get slower and costlier per turn unless the session's truncation setting drops old items. Read token usage from each response.done event rather than estimating it.
Compared with a voice stack on your own GPUs
If you run voice on your own GPUs instead, you assemble the cascade yourself: a streaming speech recogniser, a language model tuned for low time to first token, and a streaming synthesiser, usually on separate model servers. The serving characteristics differ from chat. A voice session holds its context, and therefore KV cache, for minutes while producing few tokens per second, so capacity is bounded by concurrent sessions and memory residency more than by throughput. Prefill on every turn re-reads a growing context, which is exactly where prefix caching pays off. And tail latency matters more than average, because a 2-second pause in speech feels broken in a way a slow chat reply does not.
The hosted API hides all of that, at the cost of model choice, per-minute economics you do not control, and data leaving your network. A reasonable rule: prototype on the hosted API to learn what latency and turn-taking your users need, then decide whether the volume, privacy or customisation case justifies operating your own stack.
Failure modes
- API key in the client. Always mint short-lived client secrets on the server.
- Format mismatch. 16 kHz audio declared as 24 kHz plays back fast and transcribes badly; resample or declare the right format.
- Echo self-interruption. Speaker output reaches the microphone and triggers VAD; use WebRTC echo cancellation or headsets.
- No truncate on barge-in. The model refers to sentences the user never heard.
- Silent tool calls. Slow tools leave dead air; speak a holding phrase or keep tools fast.
- Untrusted audio as instructions. Spoken or played-back audio can carry prompt injection like any other input.
- Unbounded sessions. Plan for the session limit and context growth; reconnect with a summary for long calls.
Operating it
Log every server event type with its event_id and timestamps. The metric that matters most is the gap between input_audio_buffer.speech_stopped and the first response.output_audio.delta for the same turn; track its 95th and 99th percentiles per region and transport. Also record interruptions per session, truncated milliseconds, tool latency and token usage from response.done. For abuse monitoring the WebSocket guide shows an optional OpenAI-Safety-Identifier header carrying a hashed user identifier.
Related reading
For the latency side see time-to-first-token arithmetic, streaming LLM responses and inference latency; for long sessions, KV cache sizing. On the security side, read prompt injection through audio.
What to do next
- Mint a client secret on your server and connect a browser over WebRTC with a fixed instruction set.
- Build the WebSocket loop above against a recorded audio file before touching a live microphone.
- Measure speech_stopped to first audio delta across 100 turns and record p50, p95 and p99.
- Test barge-in deliberately and confirm truncate leaves the transcript matching what was heard.
- Try server_vad and semantic_vad with real users and pick by interruption and cut-off rates, not by feel.
- Add one tool, time it, and add a holding phrase if its p95 exceeds about a second.