Media & Design

Voice (topic)

Last updated 2026-09-28

What's new

2026-09-28
  • OpenAI released cheaper GPT-6 models, and Anthropic launched Claude Opus 5.5, which is 40% cheaper and performs similarly to its more expensive Claude Fable 5.1 model.
  • Claude Opus 5.5 requires less computing power, making it more cost-effective for tasks like coding and data analysis, and it's better at communicating clearly, reducing jargon and improving readability.
  • Meta announced a new AI device small enough to carry on a keychain, and there are rumors of OpenAI releasing a new AI agent platform (a tool that automates tasks) next week.

Key points

What it is

  • Voice AI gives software the ability to speak and listen, enabling spoken conversations instead of just text.
  • It's the most natural interface for humans, reducing friction and unlocking useful actions, like a travel app guiding you to the right gate.
  • Voice AI can be a simple pipeline (speech-to-text-to-speech) or a more capable real-time voice agent (software that holds a spoken conversation).
  • It's a joint activity between the bot and user, not just a one-way pipeline.

How to use it

  • Start by planning with a wireframe (a simple flowchart) due to the complex logic involved.
  • Choose an approach: add speech-to-text and text-to-speech models, use a single speech-to-speech model, or use an out-of-the-box solution like OpenAI's Realtime API (a product for live voice conversations).
  • Configure your agent manually or use a tool like Claude Code to build and configure it for you, such as ElevenLabs (a platform for creating voice agents).
  • Feed the agent plenty of information (like books) to improve its responses, and test real calls, not just simulations.

Watch out for

  • Voice AI is error-prone, and those errors can have serious consequences, especially with embodied AI (agents with a physical presence).
  • Don't assume the latest language model intelligence carries over into voice; speed often beats cleverness in voice agents.
  • Avoid context bloat (too much information) which can lead to wrong answers; more data isn't always more helpful.
  • Building a voice agent used to mean gluing together four separate services, which can cause latency problems and other issues.

Tools named

  • OpenAI's Realtime API (product for live voice conversations), ElevenLabs (platform for creating voice agents), Claude Code (tool to build and configure agents in ElevenLabs), Gradium (trains foundation models for audio).

Lesson 1: What is Voice (topic) and why it matters

Voice, in AI development, means giving software the ability to speak and listen (spoken input and output) rather than only reading and writing text. Earlier AI voice tools worked like dictation (speech converted to text): you spoke, the system turned your words into text, the AI read that text, then converted its reply back into voice. Voice AI today often skips that relay, which is why some developers describe older systems as a simple pipeline and newer ones as more capable real-time voice agents (programs that hold a spoken conversation).

Why does it matter? First, it is the most natural interface (way for people to interact). Humans love talking, so voice reduces friction. Second, it unlocks useful agent behavior. Voice-to-action means you say what you want and the AI uses tools to get it done. Systems-to-voice means software turns live context into spoken guidance. A travel app, for example, could tell you your delayed flight connection is still possible and send you toward the right gate.

But voice carries real risk. It is error-prone, and those errors will continue for a while. The consequence of errors grows as we move from conversational bots to embodied AI (agents with a physical presence), where a wrong action matters more than a wrong answer. Poor voice design also drives user frustration, task failures, escalation to live agents, and abandoned calls. Practitioners therefore treat voice AI as a joint activity between bot and user, not just a pipeline, and test it with end-to-end benchmarks such as EVA Bench.

Sources

Lesson 2: How to use Voice (topic): step-by-step

To build a voice agent (software that talks with people), start by planning. One builder says it all starts with a wireframe (a simple flowchart) because voice systems involve so much conditional logic (rules that branch on what the caller says). Pick an approach: you can add a speech-to-text model (transcribes audio into words) on the front and a text-to-speech model on the back, which is where much of the history of speech comes from, as Google DeepMind's Tom Ouyang explains. Alternatively, use a single speech-to-speech model, like the foundation models for audio that Gradium trains. For the simplest out-of-the-box route, one course recommends OpenAI's Realtime API (a product for live voice conversations). Or build with ElevenLabs: you can configure an agent manually in the dashboard — setting the system prompt (instructions for the agent), first message, and knowledge base — or simply talk to Claude Code and let it build and configure the agent in ElevenLabs for you. A concrete example comes from a Grok bot lesson: ask it to "create a quick project where I need you to research the top five best voice AI agent providers." You can also feed the agent two to three full books of information, since most people give it only a paragraph. And if you lack preset documents, just talk to your voice tool directly.

Sources

Lesson 3: Best practices and pitfalls

A common misconception about voice agents (AI systems that handle spoken conversation) is that they always have to talk back. They don't, and that opens up more useful designs. The biggest mistake beginners make is assuming your language model's latest intelligence will carry over into voice. It mostly won't: your talking agent almost always needs "thinking turned off" so it can respond fast, which means recent LLM improvements in reasoning and instruction following don't apply. Speed beats cleverness here.

Another pitfall is architecture. Building a voice agent used to mean gluing together four separate services: speech-to-text (converting audio to words), a large language model (the reasoning engine), text-to-speech (converting words back to audio), and turn detection (deciding when the user has stopped talking). That means four dashboards, four invoices, latency (delay before a response) problems, and classic failures like mishearing a phone number or cutting you off mid-sentence. Fewer moving parts is better.

Context bloat is a quieter killer. One agency built a voice agent with a huge knowledge base, and because there was so much information, the agent couldn't find the right answer and constantly gave customers wrong information. More data is not more helpful.

Best practices: build by speaking to your computer rather than clicking through a dashboard, so you avoid forgetting to save things or misconfiguring endpoints (connection points between services). Treat voice as a genuine interface question — what should the agent sound like? — not just a text model with a voice wrapper. Finally, accept that making voice agents useful is hard even with cooperative users, so test real calls, not only simulations.

Sources