Satya Nadella just picked a public fight with the very labs his company helped build. Microsoft’s chief executive issued a blunt warning about how the biggest AI companies are rewriting the rules of data ownership, and the message landed with unusual force. In a recent blog post, Nadella accused frontier labs including OpenAI and Anthropic of constructing a rigged marketplace. They harvest public data to train massive models, then deploy their terms of service to stop anyone from learning from the outputs those models produce. The result, he argues, is a one-way transfer of value that leaves enterprises paying the bill while surrendering their proprietary knowledge.

The Distillation Double Standard

To understand the complaint, it helps to look at the technique these labs are so eager to block. Distillation is a common training method where a smaller, specialized model learns from the outputs of a larger one instead of ingesting raw internet text from scratch. A startup or research team might feed millions of prompts to a frontier model, capture the responses, and use that rich signal to train a compact model that runs locally or serves a niche need. It is cheaper and far less energy-intensive than building a foundation model from the ground up.

OpenAI and Anthropic both prohibit this practice in their terms of service. When they detect systematic distillation, they treat it as a violation. Enforcement often targets rivals, particularly AI companies based in China, who are accused of using the method to close the capability gap without matching the billions spent on original training runs.

Nadella highlighted the tension at the center of this policy. These same labs built their flagship models by crawling the open web, scraping books, forums, code repositories, and articles, all under the legal argument that training on publicly accessible data constitutes fair use. They mined the commons. Yet they now deploy legal fences to prevent anyone from treating their model outputs as public teaching material in return. You can learn from the world’s open knowledge, they seem to say, but no one may learn from us.

That asymmetry matters beyond academic debate. It establishes a hierarchy where a handful of infrastructure providers control who gets to know what. If fair use justifies their ingestion of humanity’s published intelligence, the outputs of those models begin to look like a privatized commons. Competitors and customers alike are barred from recycling that signal into new tools. The street runs only toward the largest labs.

Paying Twice for Your Own Intelligence

Nadella coined a phrase for the economic trap he sees forming: the “reverse information paradox.” In his view, enterprises are currently paying for artificial intelligence two times over, and only one of those payments shows up on an invoice.

The first cost is obvious. Companies buy API credits, subscribe to Copilot, or license chatbot services. The second cost is hidden. Every time an employee corrects a hallucinated fact, rates a response as helpful, or refines a generated report, that action produces what Nadella calls “exhaust.” It is the trail of corrections, preference data, and interaction logs generated during real work.

This exhaust is not digital debris. It often contains hard-won business logic. A logistics team might prompt a model to optimize shipping routes using internal constraints the firm spent years developing. A legal department might iteratively correct contract language until the output matches house style. A pharmaceutical research group could query clinical data in ways that reveal strategic priorities. When those interactions flow through a third-party API, the provider sees the full exchange. The model learns what good answers look like for that specific domain.

Over time, this feedback makes the provider’s general model sharper at the exact tasks the customer hired it to perform. The provider can then turn around and sell that improved capability to the customer’s competitor, or launch an adjacent product that mimics the specialized workflow. The original enterprise paid subscription fees, donated its proprietary expertise, and effectively trained its own competition.

Why Sovereign AI Is Moving From Buzzword to Strategy

ఈ నష్టానికి తార్కిక ప్రతిస్పందన ఏమిటంటే, లెర్నింగ్ లూప్ (learning loop) పై నియంత్రణను తిరిగి పొందడం. నాడెల్లా వాదన స్పష్టంగా Microsoftని ఒక విభిన్న రకమైన యజమానిగా (landlord) నిలబెట్టడానికి రూపొందించబడింది. కస్టమర్లు తమ మేధస్సును ఒక కేంద్రీకృత బ్లాక్ బాక్స్‌లోకి (centralized black box) పంపాలని పట్టుబట్టే బదులు, కస్టమర్లే తాళాలు (keys) కలిగి ఉండే వాతావరణాల కోసం ఆయన వాదిస్తున్నారు.

ఇక్కడే సోవిరీన్ AI (sovereign AI) చుట్టూ జరుగుతున్న చర్చకు ఆచరణాత్మక ప్రాధాన్యత లభిస్తుంది. ప్రైవేట్ క్లౌడ్, ఆన్-ప్రిమైస్ డేటా సెంటర్ లేదా కఠినంగా నియంత్రించబడిన వర్చువల్ నెట్‌వర్క్ లోపల మోడల్స్‌ను రన్ చేయడం అంటే, ఇంటరాక్షన్ డేటా సంస్థ యొక్క పరిధిని ఎన్నటికీ దాటదు అని అర్థం. సంస్థ ఆధునిక ఫౌండేషన్ మోడల్స్ (foundation models) నుండి ప్రయోజనం పొందుతుంది, కానీ వెయిట్స్ (weights), ప్రాంప్ట్స్ (prompts) మరియు ప్రిఫరెన్స్ డేటా (preference data) కంపెనీ నిర్వహించే సరిహద్దులోనే ఉంటాయి.

Microsoft తన Azure క్లౌడ్ వ్యాపారాన్ని సరిగ్గా ఈ ప్రైవేట్ డిప్లాయ్‌మెంట్లు మరియు రీజినల్ డేటా రెసిడెన్సీ (regional data residency) వాగ్దానం చుట్టూ నిర్మించింది. తన పోస్ట్‌లో, సంస్థలు తమ మేధో సంపత్తిని (intellectual property) తెలియకుండానే దూరంగా ఉన్న మోడల్ ప్రొవైడర్‌కు పంపకుండా, తమ స్వంత ట్రైనింగ్ మరియు ఇన్ఫరెన్స్ (training and inference) నిర్వహించుకోగలిగే ప్లాట్‌ఫామ్‌గా Microsoft ఉండాలని నాడెల్లా సూచించారు. ఈ ప్రతిపాదన సరళమైనది: శక్తివంతమైన AIని ఉపయోగించండి, కానీ మీ డేటా వ్యర్థాలను (exhaust) మీ వద్దే ఉంచుకోండి.

ఇతర వెండర్లు కూడా ఇలాంటి మాటలే చెబుతున్నారు, కానీ OpenAIతో Microsoftకి ఉన్న లోతైన భాగస్వామ్యం కారణంగా నాడెల్లా జోక్యం మరింత ప్రాధాన్యత సంతరించుకుంది. ఒక CEO కి...