getting-started-openai-realtime: The Browser Becomes the Robot Arm

A thin Cloudflare Worker keeps the keys safe, WebRTC keeps the voice loop fast, and client-side tool calls turn speech into immediate action.

10 min read · replicate/getting-started-openai-realtime

A browser window dominates a bright white page while a microphone waveform flows through a slim relay tower and into a model node. Inside the browser, a human hand flips a local switch, showing that the cloud carries the voice loop but the page performs the action.
The core idea is simple: speech travels through the cloud, but execution stays in the browser.
Key Takeaways

Most voice demos centralize everything. This one splits the job cleanly: OpenAI handles speech and tool selection over WebRTC, while the browser runs the actual function. That split is the story, because it turns the page into a command surface instead of a chat toy.

This is a real-time demo app that lets you run functions in your browser using your voice.

Replicate, Project documentation · GitHub README

The browser is the robot arm

The repo is a guide, not a product. Replicate frames it as a way to build realtime speech apps with OpenAI, Cloudflare, and Replicate, and it credits Craig Dennis at Cloudflare as the original inspiration. That matters because the stack is opinionated in a useful way: one edge worker, one browser, one realtime loop.

We tip our hat to Craig Dennis, a Developer Educator at Cloudflare who wrote the original demo that inspired this guide.

Replicate, Project documentation · Replicate guide

A thin relay, not a big backend

Everything sensitive and session-scoped stays at the edge. src/index.ts receives the browser's SDP offer at /rtc-connect, attaches the OpenAI key inside the worker, and returns the answer. The same worker can proxy /generate-image to Replicate, which keeps the frontend thin and the browser free of secret handling.

The worker relays signaling and secrets, while the browser executes tools locally and can still call Replicate for image generation.

script.js is the real orchestration layer

const tools = Object.entries(fns).map(([name, { fn, ...tool }]) => ({
  type: 'function',
  name,
  ...tool
}));

The clever part lives in the page. script.js opens getUserMedia, wires RTCPeerConnection, listens on the response data channel, and uses the fns map to resolve tool names into local JavaScript. When OpenAI emits response.function_call_arguments.done, the browser runs the function, then posts function_call_output back into the conversation.

A close-up of a browser tab showing a function registry, a data channel line, and a hand pulling one tool card from a stack while a model message arrives from the side. It explains that the browser executes the tool locally instead of forwarding the action to a server.
The page is both the interface and the executor, which is why the handoff feels instant.

Why WebRTC matters here

The transport choice is part of the product. WebRTC makes the loop feel immediate enough for speech, transcription, and tool calls to collapse into one interaction, which is why the demo feels different from older voice assistants and from many browser demos that still lean on heavier request-response plumbing.

DimensionThis repoOpenAI Realtime ConsoleGeneric WebSocket voice demo
TransportWebRTC for the realtime audio loopWebRTC for browser testingWebSocket relay to a server
Tool executionLocal functions run in the browserClient-side tool calls for experimentationUsually server-side callbacks
SecretsCloudflare Worker keeps API keys off the clientOften local or server-proxy based setupBackend usually holds the keys
Deployment shapeStatic frontend plus edge relayStandalone console for developmentApp server plus client
What it teachesA local-first action patternHow to exercise the Realtime APIBasic request-response voice plumbing

OpenAI's own guidance points browser clients toward WebRTC, and the official Realtime Console follows the same direction. Replicate's version is narrower and more instructive, because it teaches a deployment pattern, not just a demo surface.

Replicate's extra move: voice to image

A spoken command rises from a browser as a speech ribbon, condenses through a small relay, and becomes a fresh image tile inside the same page. The picture shows voice becoming a direct command surface for model generation.
The demo does not stop at speech. It can turn a spoken command into a generated image in the same interaction.

That second act is where the repo stops being just a voice assistant. A spoken prompt can trigger Replicate image generation, so the browser moves from listening to acting without a page refresh or a separate form. The point is not novelty. It is that voice becomes a command surface for multimodal output.

What this template teaches builders

The transferable lesson is clean division of labor. Keep secrets at the edge. Keep tool execution as close to the user as possible. Use WebRTC for the realtime loop, then let the browser do the work it is already good at: instant UI changes, local function calls, and fast feedback. That is the pattern this repo teaches.