Satya Nadella just picked a public fight with the very labs his company helped build. Microsoft’s chief executive issued a blunt warning about how the biggest AI companies are rewriting the rules of data ownership, and the message landed with unusual force. In a recent blog post, Nadella accused frontier labs including OpenAI and Anthropic of constructing a rigged marketplace. They harvest public data to train massive models, then deploy their terms of service to stop anyone from learning from the outputs those models produce. The result, he argues, is a one-way transfer of value that leaves enterprises paying the bill while surrendering their proprietary knowledge.
The Distillation Double Standard
To understand the complaint, it helps to look at the technique these labs are so eager to block. Distillation is a common training method where a smaller, specialized model learns from the outputs of a larger one instead of ingesting raw internet text from scratch. A startup or research team might feed millions of prompts to a frontier model, capture the responses, and use that rich signal to train a compact model that runs locally or serves a niche need. It is cheaper and far less energy-intensive than building a foundation model from the ground up.
OpenAI and Anthropic both prohibit this practice in their terms of service. When they detect systematic distillation, they treat it as a violation. Enforcement often targets rivals, particularly AI companies based in China, who are accused of using the method to close the capability gap without matching the billions spent on original training runs.
Nadella highlighted the tension at the center of this policy. These same labs built their flagship models by crawling the open web, scraping books, forums, code repositories, and articles, all under the legal argument that training on publicly accessible data constitutes fair use. They mined the commons. Yet they now deploy legal fences to prevent anyone from treating their model outputs as public teaching material in return. You can learn from the world’s open knowledge, they seem to say, but no one may learn from us.
That asymmetry matters beyond academic debate. It establishes a hierarchy where a handful of infrastructure providers control who gets to know what. If fair use justifies their ingestion of humanity’s published intelligence, the outputs of those models begin to look like a privatized commons. Competitors and customers alike are barred from recycling that signal into new tools. The street runs only toward the largest labs.
Paying Twice for Your Own Intelligence
Nadella coined a phrase for the economic trap he sees forming: the “reverse information paradox.” In his view, enterprises are currently paying for artificial intelligence two times over, and only one of those payments shows up on an invoice.
The first cost is obvious. Companies buy API credits, subscribe to Copilot, or license chatbot services. The second cost is hidden. Every time an employee corrects a hallucinated fact, rates a response as helpful, or refines a generated report, that action produces what Nadella calls “exhaust.” It is the trail of corrections, preference data, and interaction logs generated during real work.
This exhaust is not digital debris. It often contains hard-won business logic. A logistics team might prompt a model to optimize shipping routes using internal constraints the firm spent years developing. A legal department might iteratively correct contract language until the output matches house style. A pharmaceutical research group could query clinical data in ways that reveal strategic priorities. When those interactions flow through a third-party API, the provider sees the full exchange. The model learns what good answers look like for that specific domain.
Over time, this feedback makes the provider’s general model sharper at the exact tasks the customer hired it to perform. The provider can then turn around and sell that improved capability to the customer’s competitor, or launch an adjacent product that mimics the specialized workflow. The original enterprise paid subscription fees, donated its proprietary expertise, and effectively trained its own competition.
Why Sovereign AI Is Moving From Buzzword to Strategy
The logical response to this drain is to reclaim control over the learning loop. Nadella’s argument is clearly designed to position Microsoft as a different kind of landlord. Rather than insisting customers feed their intelligence into a centralized black box, he is making the case for environments where the customer keeps the keys.
This is where the conversation around sovereign AI gains practical weight. Running models inside a private cloud, an on-premise data center, or a tightly controlled virtual network means the interaction data never leaves the organization’s perimeter. The enterprise still benefits from modern foundation models, but the weights, the prompts, and the preference data remain inside a boundary the company administers.
Microsoft has built its Azure cloud business around exactly this promise of private deployments and regional data residency. In his post, Nadella signaled that Microsoft wants to be the platform where enterprises can host their own training and inference without inadvertently shipping their intellectual property to a distant model provider. The pitch is straightforward: use powerful AI, but keep your exhaust.
Other vendors are making similar noises, but Nadella’s intervention carries extra weight because of Microsoft’s deep partnership with OpenAI. For a CEO whose
