Voice assistants have long suffered from a frustrating ceiling. You could ask them the weather, set a timer, or queue up a playlist. The moment the conversation drifted toward something complex—debugging code, structuring a business deal, stress-testing a product strategy—the interaction fell apart. The assistant might hear you correctly, but it lacked the reasoning depth to think through the problem with you.
Anthropic is trying to fix that. The company has expanded voice mode in Claude from its entry-level Haiku model to its most capable systems, Opus and Sonnet. The move signals something more interesting than a simple feature rollout. Anthropic is betting that serious work can happen through spoken conversation, not just through the keyboard.
Why Voice Needs Brainpower
Previously, Claude's voice mode ran exclusively on Haiku. That model prioritizes speed and low latency. For quick questions—"What's the capital of Estonia?" or "Summarize this email?"—that trade-off made sense. Haiku fires back answers fast. But Anthropic noticed users kept pushing the feature harder than Haiku was built to handle. People tried to use voice for substantive work, and the model often came up short on the kind of extended reasoning those tasks demand.
Bringing Opus and Sonnet into the fold changes the physics of the conversation. These models are Anthropic's workhorses for difficult cognitive tasks. Opus, the flagship, handles elaborate analysis, extended reasoning chains, and creative synthesis. Sonnet sits in the middle, offering strong reasoning at greater speed than Opus. With voice layered on top, you can now talk through a messy technical architecture or a nuanced financial model and expect the AI to track the logic, spot inconsistencies, and push back with relevant questions.
Thinking Out Loud
The difference is qualitative. With Haiku voice, the interaction felt transactional. You asked; it answered. With Opus and Sonnet, the interaction becomes collaborative.
Imagine you are a product manager driving home. You dictate a half-formed idea about a new onboarding flow. Instead of receiving a generic "That's interesting" response, Claude Sonnet starts probing. It asks whether you have considered edge cases for enterprise users. It draws parallels to onboarding patterns it has seen in other contexts. By the time you reach your driveway, you have stress-tested the idea from three angles you had not considered.
Or picture an engineer troubleshooting a distributed system failure while pulling logs on a second screen. Speaking to Claude Opus, she can describe the symptoms in plain language, read error messages aloud, and talk through hypotheses without stopping to type. The model tracks the thread across multiple turns, remembers which avenues have already been ruled out, and suggests configuration changes based on the evolving diagnosis. This is not transcription with a search engine attached. It is reasoning performed in real time, through speech.
From Talking to Doing
Perhaps the most significant shift is that these advanced models can act, not merely chat. Anthropic is framing the updated voice mode as an agentic interface. Because Opus and Sonnet possess the reasoning capacity to interpret intent accurately, they can convert a spoken conversation into concrete outputs.
The example Anthropic offers is straightforward but telling. A user discusses a business idea over voice, then tells Claude to turn that rambling conversation into a structured one-page pitch. The model parses the unstructured dialogue, identifies the core value proposition, audience, and financial assumptions, and produces a formatted document. No retyping. No copying chat logs into a separate template.
This agentic layer extends into productivity tools. Anthropic is connecting the voice experience to Gmail, Slack, and Canva. That means you could dictate a project brief to Claude, ask it to generate the slide deck in Canva, drop a summary into your team's Slack channel, and draft the follow-up email in Gmail—all from the same voice session. The model becomes a coordinator, translating spoken intent into actions across separate applications.
Switching Gears Without Losing the Thread
Anthropic heeft de ervaring ook zo ontworpen dat de barrières tussen verschillende interactiemodi worden doorbroken. Gebruikers kunnen nu binnen hetzelfde gesprek schakelen tussen tekst en spraak zonder de context te verliezen. Je begint misschien met het typen van een technische vraag aan je bureau, schakelt over naar spraak terwijl je naar een vergadering loopt, en keert dan terug naar tekst om een codefragment te plakken.
Het systeem maakt het ook mogelijk om halverwege tussen verschillende modelniveaus te schakelen. Als je begint met een informele brainstormsessie met Haiku — waarbij je profiteert van de snelle reacties — kun je het gesprek opschalen naar Opus wanneer de discussie overgaat op een strikte technische uitvoering. De context wordt overgedragen. De AI onthoudt wat je drie stappen geleden hebt gezegd, ongeacht of je het hebt getypt of uitgesproken, of je nu met het snelle model of het diepgaande model sprak.
Deze vloeiendheid is belangrijk omdat echt werk zelden lineair is. Sommige problemen vereisen snelheid; andere vereisen diepgang. Het bouwen van een enkele interface waarin gebruikers tussen deze 'versnellingen' kunnen schakelen, bespaart cognitieve belasting. Het voorkomt de frictie van het openen van nieuwe tabbladen, het kopiëren van prompts of het opnieuw uitleggen van de context.
De talen van de wereld spreken
De uitbreiding is niet beperkt tot Engelstaligen. Anthropic heeft de Engelstalige bèta beëindigd en officiële ondersteuning toegevoegd voor negen extra talen:
- Europese talen: Frans, Duits, Italiaans, Spaans en Portugees.
- Aziatische talen: Hindi, Indonesisch, Japans en Koreaans.
Dit is niet zomaar een vinkje voor lokalisatie. Geavanceerde spraakassistentie met een hoog redeneervermogen in het Koreaans of Hindi verandert wie deze tools voor professioneel werk kan gebruiken. Een startup-oprichter in Jakarta kan een update voor investeerders inspreken in het Indonesisch. Een productieanalist in Turijn kan supply-chain-gegevens bevragen in het Italiaans. De cognitieve kracht van Opus wordt toegankelijk voor iedereen die natuurlijker denkt in hun moedertaal dan in het Engels.
In de bredere markt voor AI-assistenten verscherpt deze meertalige uitrol — gecombineerd met de redeneervermogens van de topmodellen — de concurrentiepositie van Claude. De meeste spraakassistenten hebben niet-Engelstalige markten historisch gezien als bijzaak behandeld, waarbij vertaling werd toegevoegd aan oppervlakkig redeneren. Anthropic lijkt erop te wedden dat wereldwijde gebruikers dezelfde diepgang van denken in hun eigen taal willen als wat Engelstaligen gewend zijn.
