Of all the debates raging about the potential downsides of artificial intelligence, one worry has caused the most hand-wringing among AI enthusiasts in Silicon Valley. Their fear is that the giant AI labs selling proprietary models are acting like Trojan horses, quietly siphoning off the most valuable asset any company possesses: its proprietary knowledge.
The concern is that as startups and enterprises use AI models from labs like OpenAI and Anthropic, these labs gain ever-increasing access to those companies’ most sensitive business information. The model makers can then use that knowledge for themselves, potentially becoming competitors to their own customers. Such warnings have been issued by venture capitalists like Jason Calacanis and Palantir CEO Alex Karp. But now, in a surprising blog post published on a Sunday, Microsoft CEO Satya Nadella has joined this crowd with a stark and nuanced argument.
Nadella warns that AI users—whom he calls the “buyers”—are paying twice. They knowingly spend for AI token usage, but they also, often obliviously, hand over valuable data in the process. “You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it!” he writes. This is a radical statement coming from the CEO of Microsoft, a company that has invested billions in OpenAI and has a deep partnership with Anthropic.
Nadella’s argument digs into the mechanics of how AI models learn. Enterprises are literally teaching the models about the nuances of their businesses. “Models learn from ‘exhaust,’ the prompts people write, the tools agents use, and especially the corrections people make when the model is wrong. Every correction is distilled into institutional know-how,” he writes. This is “the kind of knowledge a competitor could never buy,” and yet enterprises are handing it over without realizing the long-term implications.
The core of Nadella’s warning revolves around the concept of “distillation.” This is the practice of using a model’s own outputs to learn how it works and to train a new, often cheaper, model based on those insights. In February, Anthropic accused Chinese open source models of sending millions of prompts to Claude as a way to improve their own models, and urged the U.S. government to crack down on export controls. Nadella points out the hypocrisy: model makers freely train on the world’s data (scraping the internet) while imposing restrictive terms on others who want to do the same to their models. “While the great innovation that comes from model providers having fair use rights to train models on public data is needed, I find it ironic that the status quo is to then turn around and impose restrictive terms on distillation,” Nadella writes.
Nadella is particularly concerned when model makers “reserve the right to learn from customer usage and interaction data.” This is not a theoretical risk. Many enterprise customers using models from OpenAI or Anthropic implicitly agree to terms that allow the model provider to improve their models based on the prompts and feedback they receive. In effect, every time a company corrects a model’s mistake, it is donating a piece of its institutional knowledge to the model maker. Over time, this builds up a detailed picture of competitive strategies, product roadmaps, customer preferences, and internal processes.
The solution Nadella proposes is one that fits the agenda of a giant cloud provider like Microsoft. He urges companies to “retain ownership” of their data, including prompts, feedback, and corrections. To do this, he recommends building “proprietary learning environments” on the cloud—where their data is likely already stored anyway, and conveniently, could mean Microsoft’s Azure. He also advocates for building “orchestration layers” that allow companies to easily switch between AI models from different providers rather than being locked into a single vendor. Tools like AI “gateways” that enable this switching have become increasingly popular.
While Nadella never explicitly uses the words “open source” in his post, it is the obvious subtext. The only way to truly retain ownership of data and avoid vendor lock-in is to run models on your own infrastructure or on cloud instances where you control the data flow. Large companies, many of which still have some of their own data centers in addition to using the cloud, are already moving towards open source models installed on their own premises (commonly called “on-prem”). Idit Levine, founder and CEO of Solo.io, which makes networking and security software that helps enterprises manage AI systems, confirms she is seeing exactly this shift with her own customers. After experimenting with proprietary model makers, they ask themselves: “Can I take an open source model and run it on-prem? It will do almost 90% of what the big one’s doing. It will cost way less.” She adds that they understand the implications and can control the entire environment.
Solo.io’s technology was selected by the Linux Foundation to power its Agentgateway project. Her company counts enterprises like T-Mobile, ADP, and SAP as customers. Levine sees companies increasingly installing on-premise open source models and believes it is the next big wave in enterprise AI use. She is not alone. Vercel, best known as a platform for building and hosting websites, has recently added AI model-switching tools. OpenRouter, a company that helps developers route requests across different AI models, is also seeing a surge in traffic to open source models. In fact, open models accounted for 29% of all traffic routed through Vercel’s gateway last month.
This trend is driven by a combination of cost, control, and security concerns. Proprietary models from labs like OpenAI and Anthropic can be expensive to use at scale, especially when enterprises need to process millions of prompts. Open source models, such as Llama from Meta or Mistral from France, can be run on internal hardware for a fraction of the cost. More importantly, they allow enterprises to keep all data in-house, eliminating the risk of leakage to third parties. This is particularly critical for regulated industries like finance, healthcare, and defense, where data privacy laws and competitive secrecy are paramount.
Nadella’s warning also highlights a broader tension in the AI ecosystem. The same companies that are building the most advanced models are also competing with their customers in downstream markets. OpenAI has been moving into enterprise software with products like ChatGPT Enterprise and custom models for specific industries. Anthropic has its own enterprise offerings. As these labs become platforms themselves, the line between provider and competitor blurs. By urging companies to retain ownership of their data and to use orchestration layers, Nadella is effectively advising them to treat AI models as commodities rather than strategic dependencies.
The historical context here is important. The current AI boom is built on vast amounts of publicly available data scraped from the internet. Model makers have claimed fair use rights to train on this data, often over the objections of content creators and publishers. Now, these same model makers are restricting others from using their models’ outputs in similar ways. Nadella’s argument is that this asymmetry is unfair and ultimately harmful to the long-term health of the AI industry. If enterprises lose trust in proprietary models, they will either build their own or shift entirely to open source, fragmenting the market.
Microsoft’s own position is complex. The company has invested heavily in both OpenAI and Anthropic, and its Azure cloud platform hosts many of these models. At the same time, Microsoft offers its own open source models through Azure AI Foundry and has been a strong supporter of the open source community. Nadella’s blog post can be seen as a balancing act: encouraging enterprises to use the cloud (especially Azure) while warning them about the risks of proprietary lock-in. It is a message that resonates with many CIOs and CTOs who are grappling with the same dilemma.
The response from the AI labs has been muted so far. Neither OpenAI nor Anthropic have issued direct responses to Nadella’s post. However, both companies have taken steps to address data privacy concerns. OpenAI offers a “zero data retention” option for enterprise customers, and Anthropic has committed to not training on customer data without explicit permission. But these policies are not always transparent, and the details of data usage agreements can be complex.
As enterprises continue to experiment with and adopt AI, the debate over data ownership and model distillation is only going to intensify. Nadella’s warning adds significant weight to the argument that companies should be more careful about what they share with AI model providers. The trend towards open source and on-premise AI is likely to accelerate, not just because of cost savings, but because of the strategic imperative to control one’s own data. In Nadella’s own words: “In consuming intelligence, you are creating intelligence. And what you create should belong to you.”
Source: TechCrunch News