The open-source AI movement has produced genuinely impressive models. DeepSeek now processes over a third of all tokens flowing through Vercel’s AI gateway. Z.ai’s GLM-5.2 sits comfortably in the platform’s top four by volume. These are not hobbyist projects. They are production-grade systems handling real enterprise workloads at massive scale. Given these numbers, it would be easy to assume that frontier labs like Anthropic are facing an existential threat to their market position.

That assumption misses what is actually happening.

The Real Work of Frontier Models

Decagon CEO Jesse Zhang has proposed a clearer way to understand the split between proprietary and open-source AI. The competition is not a simple race to replace one another. Instead, the two categories serve different phases of a single enterprise lifecycle.

Frontier models function as the discovery layer. When a company begins exploring how AI might change an internal process, the task is almost always poorly defined. The inputs are messy. The desired outputs are vague. Success requires reasoning through ambiguity, handling edge cases the team has not yet catalogued, and adapting to prompts that change by the hour. This is prototyping and proof-of-concept work. It demands the most capable system available, cost be damned, because the alternative is a failed experiment that teaches nothing.

In this phase, an expensive frontier model is not overhead. It is the cost of market research compressed into a few API calls. Once the use case is proven and the workflow is mapped, the nature of the problem shifts. The ambiguity disappears. Inputs become standardized. Prompts stabilize. The task has become routine. At that point, many enterprises migrate the workload to a lighter, cheaper model. Open-source alternatives step in and own the production stage, while frontier models remain stationed at the frontier, handling the next wave of unknown problems.

This is not a theory about how AI should work. It is a description of how budgets are already moving.

When Workloads Move Downmarket

The migration to open source is real, and it is visible in the traffic data. DeepSeek’s surge to over one-third of token volume on Vercel’s infrastructure shows that companies are running enormous quantities of inference through cheaper models. Z.ai’s GLM-5.2 has also carved out a top-four position by handling steady, predictable traffic.

These models excel at tasks that have been tamed. Think of high-volume data extraction from standardized forms, first-pass customer support triage that routes tickets based on obvious keywords, or routine code linting and documentation generation. The prompts are templated. The error modes are understood. The business risk of a bad output is contained. When the work is defined and repetitive, the cost of inference becomes the primary concern. Running that same workload on a six-cent model instead of a premium tier makes immediate financial sense.

But volume is not revenue. The fact that open-source models dominate token counts does not mean they dominate value creation. Token volume measures activity. Token spend measures what companies are willing to pay for irreplaceable capability.

Where the Money Actually Flows

Vercel’s AI gateway data makes the economic split impossible to ignore. Despite DeepSeek’s dominance in raw token traffic, Anthropic continues to capture more than half of the total AI spend on the platform. The gap between activity and expenditure comes down to a staggering price differential.

According to OpenRouter data, Anthropic’s Opus 4.8 costs approximately $1.37 per million tokens. DeepSeek’s V4Flash costs roughly six cents for the same volume. Opus is priced about twenty-three times higher. That multiplier matters more than raw token counts ever could. A development team could move ninety percent of their inference volume to the cheaper model and still see the expensive model account for the majority of their budget.

Questa discrepanza non è un caso. Riflette la realtà secondo cui i fornitori frontier stanno vendendo qualcosa di diverso. Non vendono solo token. Vendono la capacità di ragionare su problemi che mancano di procedure stabilite. Le aziende che pagano tariffe premium non lo fanno per ignoranza. Lo fanno perché i compiti che assegnano a questi modelli sono o ad alto rischio o strutturalmente complessi. Un team legale che analizza una nuova esposizione normativa non può tollerare una citazione allucinata. Un team di prodotto che progetta un workflow agentico multi-step ha bisogno che il modello concateni correttamente la logica attraverso diversi passaggi. Il costo del fallimento in questi scenari supera di gran lunga il costo della chiamata API.

La spesa in conto capitale fluisce verso il livello in cui il valore viene ancora creato, non semplicemente eseguito.

Il mercato cresce più velocemente della migrazione

Se i modelli open-source sono molto più economici e le aziende stanno attivamente spostando i carichi di lavoro maturi verso di essi, perché la spesa per i modelli frontier non è crollata? La risposta è che il mercato totale dei compiti AI indirizzabili si sta espandendo più velocemente di quanto un singolo modello possa renderlo una commodity.

Ogni volta che un'azienda automatizza con successo un workflow prevedibile con un'alternativa open-source, accadono due cose. Primo, quel team risparmia denaro sull'esecuzione. Secondo, quelle risorse e quel talento vengono reindirizzati verso problemi adiacenti più difficili. Il lavoro di routine è ora gestito dalle macchine, il che significa che gli esseri umani possono concentrarsi sull'irregolare, sullo strategico e sull'imprevisto. Il problema appena scoperto richiede quasi sempre la profondità di ragionamento di un modello frontier.

Questo schema si ripete in tutti i settori. Una banca automatizza la revisione dei documenti utilizzando un modello economico, per poi rivolgere la sua attenzione alla costruzione di un modello di rischio dinamico che richiede un giudizio articolato. Un'azienda software automatizza la generazione di test, per poi tentare di costruire un agente di debugging autonomo che debba tracciare gli errori attraverso sistemi distribuiti. La frontiera continua ad avanzare. Non appena un compito diventa una commodity, emerge un caso d'uso più complesso che richiede capacità premium.

Molti compiti aziendali rimangono inoltre troppo sensibili per essere affidati all'attuale generazione di alternative open-source. Il supporto al triage medico, le previsioni finanziarie sotto scrutinio normativo e l'analisi della strategia esecutiva comportano rischi di perdita che rendono il costo di inferenza irrilevante rispetto ad accuratezza e affidabilità. Questi carichi di lavoro creano un livello premium duraturo. Il risultato è un'economia stabile a due livelli: un livello ad alto margine per il ragionamento complesso e la scoperta, e un livello commodity ad alto volume per l'esecuzione della produzione di routine.

Cosa significa questo per gli acquirenti enterprise

La conclusione pratica è che la selezione del modello dovrebbe seguire la maturità del lavoro, non l'ideologia. Costruisci e valida nuove applicazioni AI sui modelli frontier più capaci a cui puoi accedere. Paga il premio durante la fase di scoperta. Costa meno che costruire su un modello limitato, non riuscire a dimostrarne il valore e abbandonare il progetto. Una volta noti gli input, gli output e le modalità di errore, ottimizza aggressivamente. Sposta il carico di lavoro stabile su un'alternativa open-source e cattura il risparmio sui costi.

Cercare di forzare ogni compito in un unico livello è una ricetta per sprecare capitale o perdere capacità. Le aziende che navigheranno correttamente in questo scenario utilizzeranno architetture ibride per impostazione predefinita, non come un compromesso.

La storia qui non è che l'open source stia perdendo, o che i laboratori frontier siano invincibili. È che entrambi i livelli stanno crescendo, ma crescono in direzioni diverse. I modelli open-source stanno assorbendo il mondo noto dei compiti AI. I modelli frontier stanno rivendicando il territorio dell'ignoto. Per il futuro prevedibile, questo è un accordo confortevole per entrambi.