Voice in AI used to mean dictation, or an assistant you spoke to one turn at a time. OpenAI's GPT-Live-1, released in the API on September 10, 2026, shows a different shape: a voice model that holds the conversation while a separate model does the work. That split changes what voice products can do, what they cost, and where they break.

What actually changed

Older voice assistantsNewer full-duplex voice
Wait for you to finish, then respondListen while speaking, so you can interrupt
Pause awkwardly or cut you off mid-thoughtHandle pauses and corrections more naturally
One system hears, thinks and actsA voice model talks while a backend model reasons and uses tools
Answer questionsStart and steer tasks while the conversation continues

GPT-Live-1 in concrete terms:

  • It listens and speaks at the same time, and hands reasoning and tool use to a backend model.
  • That backend can be an OpenAI model or, with client delegation, any model, agent or service you run.
  • It takes audio and text only, with no image input.
  • It connects to phone lines through SIP and partners like Twilio and LiveKit.
  • It costs $0.05 per minute, billed per second, with backend usage billed separately.

Early users include Speak, Intercom's Fin, Yelp Host and Cognition, whose Devin Voice lets developers talk through and hand off coding work. OpenAI reports that in Speak's testing, false interruptions during learners' thinking pauses dropped by nearly 80% compared with turn-based systems. That's the difference between a tutor that lets you think and one that jumps in.

The voice is the front desk, not the brain

This is the least obvious change, and it matters most for anyone building.

The voice model isn't where the knowledge lives. GPT-Live-1's own knowledge cutoff is July 31, 2025. Current information, business rules and tools belong in the backend. OpenAI's docs say to keep speaking style in the voice prompt and procedures in the backend prompt.

You can change the brain without changing the voice. Because the backend is chosen separately, a team can switch reasoning models, or run its existing text agent behind voice, without rebuilding the conversation layer.

"Stop talking" isn't "cancel." In GPT-Live, interrupting speech doesn't stop backend work. If you cut the assistant off mid-sentence, a booking it started may still go through. OpenAI's docs say the voice model must not claim an action finished before the backend confirms it, and that your application owns permissions and confirmations.

The design lesson: keep what's being said separate from what's being done, and confirm actions in words the user can't misread.

Talk to give context, read to review

Voice works best in one direction.

  • Input: in a Stanford and University of Washington study, speech entry was about 3x faster than typing on a phone keyboard, with fewer errors.
  • Output: adults silently read nonfiction at about 238 words per minute on average, and reading lets you skim. Audio makes you listen in order.

So the productive pattern is voice in, screen out. Talk through the messy version of what you want, including the context you'd never bother typing, then review the result as text.

You don't need an API to try it:

  • Dictate prompts: Ctrl+Shift+D in the ChatGPT desktop app, Caps Lock with quick entry in Claude for Mac once you turn it on, or an app like Wispr Flow in any text field.
  • Steer work by voice: ChatGPT Voice in the desktop app can start tasks, check progress and change direction. In Claude's newer experience, voice mode works in conversations where Claude is carrying out a task.

Where voice lands first

Use caseWhy voice fitsWatch for
Support and phone linesCallers already talk. No phone tree, and interruptions workClear handoff to a person, and confirmation before account changes
Intake and schedulingMissed calls become booked appointmentsWrong dates or names repeated back without checking
Language learning and tutoringPauses and pronunciation are part of the lessonOver-correcting, or letting mistakes slide
Practice conversationsInterviews, sales calls and hard conversations need a live partnerFeedback that sounds confident but isn't specific
Coding and technical workExplaining intent out loud, then reviewing the diff on screenSpoken changes applied without reading them
Hands-busy workField service, cooking, commuting, lab workNoise, and actions taken without a screen to confirm

The costs people miss

  • Silence is billable. GPT-Live bills active session time, including when nobody's speaking and when the backend is still working. A slow tool call costs voice minutes as well as backend tokens.
  • The math is small but constant. $0.05 per minute is $3 per hour of open session before any backend costs.
  • Starting a session isn't free. Creating a WebRTC session bills 15 seconds up front, which is then credited against the session. Opening sessions before users are ready adds up.
  • Speed saves money. OpenAI's guidance is to load context before the call starts and make backend calls faster, since finishing a minute sooner saves that minute.

Risks that grow with voice

  • Actions without a clear yes. Spoken commands are easy to mishear and hard to review. Anything consequential should get an explicit confirmation.
  • Recording in shared spaces. A voice assistant hears everyone nearby, not just you.
  • Screens you share by accident. Desktop voice features that "take a look at this" can capture a whole window, including text outside the visible area.
  • Voice is no longer proof of identity. When software can sound natural on a live call, a familiar voice alone shouldn't authorize anything.

The bigger picture

  • The interface becomes the conversation. Instead of opening apps and menus, you describe the goal and the agent coordinates tools behind the voice.
  • Products compete on the backend. When voice layers are available per minute, the difference is the tools, data, permissions and workflows behind them.
  • New devices get simpler. A microphone, a speaker and a network connection can front a capable agent.
  • Typing doesn't disappear. Precise edits, code, quiet offices and anything you need to review line by line still belong on a screen.

Voice won't replace every interface. It's becoming the fastest way to hand work to an agent, and screens are becoming where you check it.

Go deeper: Stop Building Your AI System Inside One Tool · The AI Tool Map · ChatGPT, End to End