Satya Nadella just picked a public fight with the very labs his company helped build. Microsoft’s chief executive issued a blunt warning about how the biggest AI companies are rewriting the rules of data ownership, and the message landed with unusual force. In a recent blog post, Nadella accused frontier labs including OpenAI and Anthropic of constructing a rigged marketplace. They harvest public data to train massive models, then deploy their terms of service to stop anyone from learning from the outputs those models produce. The result, he argues, is a one-way transfer of value that leaves enterprises paying the bill while surrendering their proprietary knowledge.
The Distillation Double Standard
To understand the complaint, it helps to look at the technique these labs are so eager to block. Distillation is a common training method where a smaller, specialized model learns from the outputs of a larger one instead of ingesting raw internet text from scratch. A startup or research team might feed millions of prompts to a frontier model, capture the responses, and use that rich signal to train a compact model that runs locally or serves a niche need. It is cheaper and far less energy-intensive than building a foundation model from the ground up.
OpenAI and Anthropic both prohibit this practice in their terms of service. When they detect systematic distillation, they treat it as a violation. Enforcement often targets rivals, particularly AI companies based in China, who are accused of using the method to close the capability gap without matching the billions spent on original training runs.
Nadella highlighted the tension at the center of this policy. These same labs built their flagship models by crawling the open web, scraping books, forums, code repositories, and articles, all under the legal argument that training on publicly accessible data constitutes fair use. They mined the commons. Yet they now deploy legal fences to prevent anyone from treating their model outputs as public teaching material in return. You can learn from the world’s open knowledge, they seem to say, but no one may learn from us.
That asymmetry matters beyond academic debate. It establishes a hierarchy where a handful of infrastructure providers control who gets to know what. If fair use justifies their ingestion of humanity’s published intelligence, the outputs of those models begin to look like a privatized commons. Competitors and customers alike are barred from recycling that signal into new tools. The street runs only toward the largest labs.
Paying Twice for Your Own Intelligence
Nadella coined a phrase for the economic trap he sees forming: the “reverse information paradox.” In his view, enterprises are currently paying for artificial intelligence two times over, and only one of those payments shows up on an invoice.
The first cost is obvious. Companies buy API credits, subscribe to Copilot, or license chatbot services. The second cost is hidden. Every time an employee corrects a hallucinated fact, rates a response as helpful, or refines a generated report, that action produces what Nadella calls “exhaust.” It is the trail of corrections, preference data, and interaction logs generated during real work.
This exhaust is not digital debris. It often contains hard-won business logic. A logistics team might prompt a model to optimize shipping routes using internal constraints the firm spent years developing. A legal department might iteratively correct contract language until the output matches house style. A pharmaceutical research group could query clinical data in ways that reveal strategic priorities. When those interactions flow through a third-party API, the provider sees the full exchange. The model learns what good answers look like for that specific domain.
Over time, this feedback makes the provider’s general model sharper at the exact tasks the customer hired it to perform. The provider can then turn around and sell that improved capability to the customer’s competitor, or launch an adjacent product that mimics the specialized workflow. The original enterprise paid subscription fees, donated its proprietary expertise, and effectively trained its own competition.
Why Sovereign AI Is Moving From Buzzword to Strategy
Phản ứng hợp lý trước sự thất thoát này là giành lại quyền kiểm soát vòng lặp học tập. Lập luận của Nadella rõ ràng được thiết kế để định vị Microsoft như một kiểu "chủ cho thuê" khác biệt. Thay vì yêu cầu khách hàng nạp trí tuệ của họ vào một hộp đen tập trung, ông đang ủng hộ các môi trường nơi khách hàng nắm giữ chìa khóa.
Đây chính là lúc các cuộc thảo luận xoay quanh AI chủ quyền (sovereign AI) trở nên có sức nặng thực tiễn. Việc vận hành các mô hình bên trong một đám mây riêng (private cloud), một trung tâm dữ liệu tại chỗ (on-premise), hoặc một mạng ảo được kiểm soát chặt chẽ có nghĩa là dữ liệu tương tác sẽ không bao giờ rời khỏi phạm vi của tổ chức. Doanh nghiệp vẫn được hưởng lợi từ các mô hình nền tảng hiện đại, nhưng các trọng số (weights), các câu lệnh (prompts) và dữ liệu sở thích vẫn nằm trong ranh giới mà công ty quản lý.
Microsoft đã xây dựng mảng kinh doanh đám mây Azure xoay quanh chính lời hứa về việc triển khai riêng tư và lưu trú dữ liệu theo khu vực (regional data residency). Trong bài đăng của mình, Nadella đã ra tín hiệu rằng Microsoft muốn trở thành nền tảng nơi các doanh nghiệp có thể tự lưu trữ việc huấn luyện và suy luận (inference) mà không vô tình gửi tài sản trí tuệ của họ đến một nhà cung cấp mô hình ở xa. Lời chào hàng rất thẳng thắn: sử dụng AI mạnh mẽ, nhưng hãy giữ lại các dữ liệu phát sinh của chính mình.
Các nhà cung cấp khác cũng đang đưa ra những tuyên bố tương tự, nhưng sự can thiệp của Nadella mang sức nặng lớn hơn nhờ mối quan hệ đối tác sâu rộng của Microsoft với OpenAI. Đối với một CEO mà
