2026年6月中旬の数字が出揃い、紛れもない状況が浮き彫りになった。中国のAIモデルは、トークン量において10週連続で世界トップの座を維持している。これは単一のベンチマークによる一時的な見出しでも、一つのバイラルなリリースに伴う一時的な急増でもない。人工知能の実際の作業がどこで行われているかを明らかにする、数週間にわたる持続的なパターンである。
世界の週間総ボリューム46.7兆トークンのうち、中国のモデルは18.81兆トークンを処理した。アメリカのモデルは5.76兆トークンであった。これは中国が46%、米国が13%に相当する。トークン量は、実社会におけるAI導入を示す最も具体的な指標である。すべてのトークンは、レビューされたコードの一行、回答された顧客の問い合わせ、要約された文書、説明された画像、実行された推論チェーンといった、計算労働の実際の単位を表している。世界のAI労働のほぼ半分が中国で構築されたモデルを通じて行われているということは、グローバルなAI経済の根本的な再分配が起きていることを意味している。
企業への移行は急速に進んだ
米国企業内でのシフトは急激であった。米国企業による中国モデルのトークン消費量は、2025年の4.5%から2026年には46%へと上昇した。2026年2月以降、そのシェアは30%を上回った状態が続いている。これら2つのデータポイントは、明確なストーリーを物語っている。これは、一部の好奇心旺盛なエンジニアによる初期の実験ではない。今年初めに閾値を超え、その後後戻りすることのない、広範で構造的な移行である。
米国企業は当初、中国のモデルをバックアップの選択肢や好奇心の対象として扱っていた。しかし、チームが内部的なコスト比較を開始した。その結果、切り替えても日常的なタスクにおける精度を犠牲にする必要がないことが判明した。それが明らかになると、調達の決定は迅速に進んだ。企業シェアが46%であるということは、米国内のAI計算予算のほぼ半分が、今や太平洋の向こう側で開発されたアーキテクチャに流れていることを意味する。サンフランシスコとベイエリアを重力の中心として長年扱ってきた業界にとって、これは驚くべき再編である。
実在する企業、実在する節約
有名企業の具体的な決定は、この傾向がいかに深いものかを示している。暗号資産取引所のCoinbaseは、エンジニア向けにGLM-5.2とKimi K2.7を選択した。これは研究室に埋もれたパイロットプログラムではなかった。これらのモデルは、コード補完、技術文書、デバッグ支援、内部ツールといった、実際の開発ワークフローを支えている。Coinbaseは厳格な稼働時間要件とセキュリティ体制を維持している。同社のエンジニアリングチームは、わずかな利益のために外国のモデルを採用したのではない。モデルが企業の信頼性基準を満たしつつ、優れた経済性を提供したからこそ採用したのである。
また、Anthropic ClaudeからDeepSeek-V4に移行した米国のスタートアップ、Lindyの例もある。その結果、コストを95%削減し、数百万ドルを節約することに成功した。スタートアップは限られたランウェイで運営されている。推論コストの95%削減は、18ヶ月で資金が底をつくか、それとも数年間にわたって余裕を持って規模を拡大できるかの分かれ目になり得る。しかし、Lindyの動きはより広範なことも示唆している。つまり、DeepSeek-V4が本番環境においてClaudeに取って代わるのに十分な性能を備えていたということだ。これは予算重視の低スペックへのダウングレードではなく、実用性を維持しながらコストを劇的に抑えるための入れ替えであった。
これら2つの例は、企業のスペクトルの両極端に位置している。Coinbaseは、コンプライアンスチームとレガシーなインフラを備えた上場テック企業である。Lindyは、賢明なAI経済性に生存を賭けている初期段階の企業である。両者が同じ結論に達した。その事実は、我々に何かを伝えているはずだ。
実践における「効率性」の真の意味
効率性がこれらの選択を後押ししているが、その意味については具体的に定義すべきである。それは単にダッシュボード上のAPI価格が安くなったということではない。中国の研究所は、各計算サイクルから大幅に多くのパフォーマンスを引き出す推論アーキテクチャを生み出してきた。より優れた量子化、最適化されたアテンション・メカニズム、蒸留されたモデルのバリアント、そしてハードウェアを意識したサービング・スタック。これらが組み合わさることで、トークンあたりのコストを押し下げているのである。
Why does that matter so much? Because token volume is not static. When a company builds a successful AI feature, usage tends to compound. If your application generates ten times as many tokens next quarter because customers love it, your infrastructure bill scales with that growth. A model that is merely “cheaper” helps. A model that is ninety-five percent cheaper changes your unit economics entirely. It determines whether your AI productline is profitable or a burn center. It lets startups compete with incumbents and lets incumbents protect their margins while shipping more AI capabilities.
American firms are waking up to this math. They are discovering that many production tasks—routing tickets, drafting emails, parsing logs, generating test cases, summarizing meetings—do not require the most expensive frontier model on the market. They require a model that is good enough at a cost structure that makes the business model work. Chinese providers have stepped into that gap aggressively.
Reading the Ten-Week Streak
Ten weeks at the top of global token volume is a long time in AI. A single week could be an anomaly. Ten weeks is a trend with momentum. It suggests that Chinese models have moved past the “evaluation” phase inside global enterprises and into the “default” phase. Engineers are not just testing them; they are building on them. Product managers are allocating budget to them. The infrastructure is being integrated into continuous deployment pipelines and customer-facing systems.
The February 2026 tipping point makes sense in hindsight. Models like DeepSeek-V4, GLM-5.2, and Kimi K2.7 had been available for some time, but early 2026 appears to be when American teams gained enough production experience to trust them at scale. Once trust crossed a critical threshold, the enterprise share jumped above 30% and kept climbing. By mid-June, Chinese models were handling nearly four times the token volume of US models globally.
The Takeaway for Builders
If you are building products or running engineering teams, the practical lesson is straightforward. Stop equating model quality with zip code. The best architecture for your specific workload might not come from a Bay Area provider. Run your own cost-per-task benchmarks on real data. Measure latency, accuracy, and price together. Consider what happens to your budget when usage scales by 10x or 100x.
The global AI infrastructure layer is globalizing fast. Cost efficiency is now the primary engine reshaping enterprise adoption, and China’s labs have spent the last year optimizing exactly for that pressure. The result is a market where 46% of the world’s AI tokens flow through Chinese models, and American companies now account for nearly half of that enterprise usage themselves. The geography of AI’s heavy lifting has shifted. The numbers do not lie.
Source: China Dominates Global Token Volume
Optional learning community: GyaanSetu AI on Telegram
