Voice assistants have long suffered from a frustrating ceiling. You could ask them the weather, set a timer, or queue up a playlist. The moment the conversation drifted toward something complex—debugging code, structuring a business deal, stress-testing a product strategy—the interaction fell apart. The assistant might hear you correctly, but it lacked the reasoning depth to think through the problem with you.

Anthropic is trying to fix that. The company has expanded voice mode in Claude from its entry-level Haiku model to its most capable systems, Opus and Sonnet. The move signals something more interesting than a simple feature rollout. Anthropic is betting that serious work can happen through spoken conversation, not just through the keyboard.

Why Voice Needs Brainpower

Previously, Claude's voice mode ran exclusively on Haiku. That model prioritizes speed and low latency. For quick questions—"What's the capital of Estonia?" or "Summarize this email?"—that trade-off made sense. Haiku fires back answers fast. But Anthropic noticed users kept pushing the feature harder than Haiku was built to handle. People tried to use voice for substantive work, and the model often came up short on the kind of extended reasoning those tasks demand.

Bringing Opus and Sonnet into the fold changes the physics of the conversation. These models are Anthropic's workhorses for difficult cognitive tasks. Opus, the flagship, handles elaborate analysis, extended reasoning chains, and creative synthesis. Sonnet sits in the middle, offering strong reasoning at greater speed than Opus. With voice layered on top, you can now talk through a messy technical architecture or a nuanced financial model and expect the AI to track the logic, spot inconsistencies, and push back with relevant questions.

Thinking Out Loud

The difference is qualitative. With Haiku voice, the interaction felt transactional. You asked; it answered. With Opus and Sonnet, the interaction becomes collaborative.

Imagine you are a product manager driving home. You dictate a half-formed idea about a new onboarding flow. Instead of receiving a generic "That's interesting" response, Claude Sonnet starts probing. It asks whether you have considered edge cases for enterprise users. It draws parallels to onboarding patterns it has seen in other contexts. By the time you reach your driveway, you have stress-tested the idea from three angles you had not considered.

Or picture an engineer troubleshooting a distributed system failure while pulling logs on a second screen. Speaking to Claude Opus, she can describe the symptoms in plain language, read error messages aloud, and talk through hypotheses without stopping to type. The model tracks the thread across multiple turns, remembers which avenues have already been ruled out, and suggests configuration changes based on the evolving diagnosis. This is not transcription with a search engine attached. It is reasoning performed in real time, through speech.

From Talking to Doing

Perhaps the most significant shift is that these advanced models can act, not merely chat. Anthropic is framing the updated voice mode as an agentic interface. Because Opus and Sonnet possess the reasoning capacity to interpret intent accurately, they can convert a spoken conversation into concrete outputs.

The example Anthropic offers is straightforward but telling. A user discusses a business idea over voice, then tells Claude to turn that rambling conversation into a structured one-page pitch. The model parses the unstructured dialogue, identifies the core value proposition, audience, and financial assumptions, and produces a formatted document. No retyping. No copying chat logs into a separate template.

This agentic layer extends into productivity tools. Anthropic is connecting the voice experience to Gmail, Slack, and Canva. That means you could dictate a project brief to Claude, ask it to generate the slide deck in Canva, drop a summary into your team's Slack channel, and draft the follow-up email in Gmail—all from the same voice session. The model becomes a coordinator, translating spoken intent into actions across separate applications.

Switching Gears Without Losing the Thread

Anthropic has also designed the experience to break down barriers between different modes of interaction. Users can now jump between text and voice within the same conversation without losing context. You might start typing a technical query at your desk, switch to voice while walking to a meeting, then return to text to paste a code snippet.

The system also allows switching between model tiers midstream. If you begin a casual brainstorming session with Haiku—taking advantage of its snappy responses—you can escalate the conversation to Opus when the discussion turns to rigorous technical execution. The context transfers over. The AI remembers what you said three turns ago, regardless of whether you typed it or spoke it, or whether you were talking to the fast model or the deep one.

This fluidity matters because real work is rarely linear. Some problems need speed; others need depth. Building a single interface where users can move between those gears saves cognitive overhead. It prevents the friction of opening new tabs, copying prompts, or re-explaining context.

Speaking the World's Languages

The expansion is not limited to English speakers. Anthropic has ended the English-only beta, adding official support for nine additional languages:

  • European languages: French, German, Italian, Spanish, and Portuguese.
  • Asian languages: Hindi, Indonesian, Japanese, and Korean.

This is not merely a localization checkbox. High-reasoning voice assistance in Korean or Hindi changes who can use these tools for professional work. A startup founder in Jakarta can voice-draft an investor update in Indonesian. A manufacturing analyst in Turin can interrogate supply-chain data in Italian. The cognitive power of Opus becomes accessible to anyone who thinks more naturally in their native tongue than in English.

In the broader market for AI assistants, this multilingual rollout—paired with the reasoning capabilities of the top-tier models—sharpens Claude's competitive edge. Most voice assistants have historically treated non-English markets as afterthoughts, layering translation onto shallow reasoning. Anthropic seems to be betting that global users want the same depth of thought in their own languages that English speakers have come to expect.

The Real Take