जून २०२६ च्या मधल्या काळातील आकडेवारी समोर आली आहे आणि ती एक स्पष्ट चित्र मांडते. चिनी AI मॉडेल्सनी सलग दहा आठवडे टोकन व्हॉल्यूममध्ये जागतिक स्तरावर अव्वल स्थान राखले आहे. ही केवळ एखाद्या बेंचमार्कमधील तात्पुरती बातमी किंवा एखाद्या व्हायरल रिलीजमुळे आलेली तात्पुरती वाढ नाही. ही एक सातत्यपूर्ण, आठवड्यांच्या कालावधीची अशी पद्धत आहे जी कृत्रिम बुद्धिमत्तेचे (AI) प्रत्यक्ष काम कोठे होत आहे हे दर्शवते.

एकूण जागतिक साप्ताहिक व्हॉल्यूम ४६.७ ट्रिलियन टोकन्सपैकी, चिनी मॉडेल्सनी १८.८१ ट्रिलियन टोकन्सवर प्रक्रिया केली. अमेरिकन मॉडेल्सनी ५.७६ ट्रिलियन टोकन्स हाताळले. याचा अर्थ चीनसाठी ४६% आणि युनायटेड स्टेट्ससाठी १३% असा होतो. टोकन व्हॉल्यूम हे प्रत्यक्ष वापरामध्ये AI च्या अवलंबनाचे (adoption) सर्वात ठोस मोजमाप आहे. प्रत्येक टोकन हे संगणकीय श्रमाचे (computational labor) एक वास्तविक युनिट दर्शवते—जसे की तपासलेली कोडची एक ओळ, उत्तर दिलेला ग्राहकाचा प्रश्न, सारांशित केलेला दस्तऐवज, वर्णन केलेले चित्र किंवा कार्यान्वित केलेली रिझनिंग चेन. जेव्हा जगातील जवळपास अर्धे AI श्रम चीनमध्ये तयार केलेल्या मॉडेल्सद्वारे चालवले जातात, तेव्हा आपण जागतिक AI अर्थव्यवस्थेच्या मूलभूत पुनर्वितरणाकडे (redistribution) पाहत असतो.

एंटरप्राइझ स्थलांतर वेगाने झाले

अमेरिकन कंपन्यांमधील हा बदल अत्यंत तीव्र आहे. अमेरिकन एंटरप्राइझमधील चिनी मॉडेल्सचा टोकन वापर २०२५ मधील ४.५% वरून २०२६ मध्ये ४६% पर्यंत वाढला आहे. फेब्रुवारी २०२६ पासून, हा वाटा ३०% च्या वर राहिला आहे. हे दोन डेटा पॉइंट्स मिळून एक स्पष्ट कथा सांगतात. हे काही मोजक्या जिज्ञासू इंजिनिअर्सद्वारे केलेले सुरुवातीचे प्रयोग नाहीत. हे एक व्यापक, संरचनात्मक स्थलांतर आहे ज्याने या वर्षाच्या सुरुवातीला एक मर्यादा ओलांडली आणि त्यानंतर मागे वळून पाहिले नाही.

अमेरिकन व्यवसायांनी सुरुवातीला चिनी मॉडेल्सकडे बॅकअप पर्याय किंवा केवळ कुतूहल म्हणून पाहिले. त्यानंतर टीम्सनी अंतर्गत खर्च तुलना (cost comparisons) करण्यास सुरुवात केली. त्यांना असे आढळले की, स्विच केल्यामुळे नियमित कामांमधील अचूकतेशी तडजोड करण्याची गरज नाही. एकदा हे स्पष्ट झाले की, खरेदीचे (procurement) निर्णय वेगाने घेतले गेले. ४६% एंटरप्राइझ वाटा याचा अर्थ असा की अमेरिकन कंपन्यांमधील AI कॉम्प्युट बजेटचा जवळपास अर्धा भाग आता पॅसिफिकच्या पलीकडे विकसित केलेल्या आर्किटेक्चरकडे वळत आहे. ज्या उद्योगाने सॅन फ्रान्सिस्को आणि बे एरियाला आपले केंद्र मानले आहे, त्यांच्यासाठी हे एक लक्षणीय पुनर्रचना (realignment) आहे.

वास्तविक कंपन्या, वास्तविक बचत

नामांकित कंपन्यांचे ठोस निर्णय हे हे बदल किती खोलवर आहेत हे दर्शवतात. क्रिप्टोकरन्सी एक्सचेंज, Coinbase ने त्यांच्या इंजिनिअर्ससाठी GLM-5.2 आणि Kimi K2.7 निवडले. हा केवळ रिसर्च लॅबमध्ये दडलेला एखादा पायलट प्रोग्राम नव्हता. ही मॉडेल्स प्रत्यक्ष डेव्हलपर वर्कफ्लो—कोड पूर्ण करणे (code completion), तांत्रिक दस्तऐवजीकरण (technical documentation), डीबगिंग असिस्टन्स आणि अंतर्गत टूल्स—यांना शक्ती देतात. Coinbase च्या अपटाइम आणि सुरक्षा मानकांची (security postures) कडक आवश्यकता आहेत. त्यांच्या इंजिनिअरिंग टीमने केवळ किरकोळ फायद्यासाठी परदेशी मॉडेल्सचा अवलंब केला नाही. त्यांनी ते स्वीकारले कारण या मॉडेल्सनी एंटरप्राइझच्या विश्वासार्हतेचे निकष पूर्ण केले आणि सोबतच उत्कृष्ट आर्थिक फायदा (economics) दिला.

त्यानंतर Lindy चा उल्लेख येतो, ही एक अमेरिकन स्टार्टअप आहे जिने Anthropic Claude कडून DeepSeek-V4 कडे स्थलांतर केले. याचा परिणाम ९५% खर्च कपात आणि लाखो डॉलर्सची बचत असा झाला. स्टार्टअप्स अत्यंत मर्यादित निधीवर (runways) काम करतात. इन्फरन्स खर्चात ९५% कपात करणे म्हणजे अठरा महिन्यांत रोख रक्कम संपण्याऐवजी अनेक वर्षे सहज विस्तारणे (scaling) असा असू शकतो. परंतु Lindy च्या या निर्णयाने आणखी काही व्यापक संकेत दिले आहेत: DeepSeek-V4 ने प्रोडक्शन एन्व्हायरमेंटमध्ये Claude ची जागा घेण्यासाठी पुरेशी कामगिरी केली. हे केवळ बजेट पर्यायाकडे केलेले downgrade नव्हते. हे एक असे बदल होते ज्याने उपयुक्तता कायम ठेवली आणि खर्च मात्र पूर्णपणे कमी केला.

ही दोन उदाहरणे कॉर्पोरेट क्षेत्राच्या दोन टोकांवर आहेत. Coinbase ही कंप्लायन्स टीम्स आणि जुन्या पायाभूत सुविधा (legacy infrastructure) असलेली एक सार्वजनिकरित्या सूचीबद्ध टेक दिग्गज कंपनी आहे. Lindy ही एक सुरुवातीच्या टप्प्यातील संस्था आहे जी आपल्या अस्तित्वासाठी स्मार्ट AI अर्थशास्त्रावर (AI economics) अवलंबून आहे. दोन्ही एकाच निष्कर्षापर्यंत पोहोचल्या. त्यातून आपल्याला काहीतरी समजले पाहिजे.

कार्यक्षमता (Efficiency) म्हणजे प्रत्यक्षात काय?

कार्यक्षमता या निवडींना चालना देत आहे, परंतु त्याचा नेमका अर्थ काय आहे याबद्दल आपण स्पष्ट असले पाहिजे. हे केवळ डॅशबोर्डवरील स्वस्त API किमतींचा विषय नाही. चिनी लॅब्सनी अशा इन्फरन्स आर्किटेक्चरची निर्मिती केली आहे जे प्रत्येक कॉम्प्युट सायकलमधून लक्षणीयरीत्या अधिक कामगिरी मिळवून देतात. उत्तम क्वांटायझेशन (quantization), ऑप्टिमाइझ्ड अटेंशन मेकॅनिझम (optimized attention mechanisms), डिस्टिल्ड मॉडेल व्हेरिएंट्स आणि हार्डवेअर-अवेअर सर्व्हिंग स्टॅक्स या सर्वांच्या एकत्रित प्रयत्नांमुळे प्रति टोकन खर्च कमी होत आहे.

Why does that matter so much? Because token volume is not static. When a company builds a successful AI feature, usage tends to compound. If your application generates ten times as many tokens next quarter because customers love it, your infrastructure bill scales with that growth. A model that is merely “cheaper” helps. A model that is ninety-five percent cheaper changes your unit economics entirely. It determines whether your AI productline is profitable or a burn center. It lets startups compete with incumbents and lets incumbents protect their margins while shipping more AI capabilities.

American firms are waking up to this math. They are discovering that many production tasks—routing tickets, drafting emails, parsing logs, generating test cases, summarizing meetings—do not require the most expensive frontier model on the market. They require a model that is good enough at a cost structure that makes the business model work. Chinese providers have stepped into that gap aggressively.

Reading the Ten-Week Streak

Ten weeks at the top of global token volume is a long time in AI. A single week could be an anomaly. Ten weeks is a trend with momentum. It suggests that Chinese models have moved past the “evaluation” phase inside global enterprises and into the “default” phase. Engineers are not just testing them; they are building on them. Product managers are allocating budget to them. The infrastructure is being integrated into continuous deployment pipelines and customer-facing systems.

The February 2026 tipping point makes sense in hindsight. Models like DeepSeek-V4, GLM-5.2, and Kimi K2.7 had been available for some time, but early 2026 appears to be when American teams gained enough production experience to trust them at scale. Once trust crossed a critical threshold, the enterprise share jumped above 30% and kept climbing. By mid-June, Chinese models were handling nearly four times the token volume of US models globally.

The Takeaway for Builders

If you are building products or running engineering teams, the practical lesson is straightforward. Stop equating model quality with zip code. The best architecture for your specific workload might not come from a Bay Area provider. Run your own cost-per-task benchmarks on real data. Measure latency, accuracy, and price together. Consider what happens to your budget when usage scales by 10x or 100x.

The global AI infrastructure layer is globalizing fast. Cost efficiency is now the primary engine reshaping enterprise adoption, and China’s labs have spent the last year optimizing exactly for that pressure. The result is a market where 46% of the world’s AI tokens flow through Chinese models, and American companies now account for nearly half of that enterprise usage themselves. The geography of AI’s heavy lifting has shifted. The numbers do not lie.

Source: China Dominates Global Token Volume

Optional learning community: GyaanSetu AI on Telegram