This assertion is not mere conjecture but is rooted in a meticulous examination of the contractual fine print governing the use of leading AI services. Companies like OpenAI and Anthropic, the dominant players in the generative AI space, frequently assure their commercial clientele that their proprietary data will not be used for model training. This promise is a cornerstone of trust, vital for encouraging enterprises to integrate powerful, yet data-hungry, AI systems into their core operations. However, as one recent experience revealed—after moving my own company off Anthropic following a critical Supply-Chain Risk designation—a closer reading of the actual agreements exposes a loophole capacious enough to encompass the entire intellectual property of an unsuspecting enterprise.
Consider the language embedded within the OpenAI Services Agreement. It states: "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use." On the surface, this appears to be a reasonable and protective clause, designed to safeguard client confidentiality and competitive advantage. The devil, however, lies in the definitions. "Customer Content" is defined as both "Input" and "Output." "Input" is straightforward enough: it’s the data or query a customer sends to the model. "Output" is what the model returns based on that Input. Yet, this seemingly clear distinction is far more ambiguous and strategically vague than it initially appears, particularly when confronted with the sophisticated inner workings of modern reasoning models.
The Hidden Tokens: Unveiling the AI’s "Scratchpad"
The evolution of large language models has been rapid and transformative. Early iterations of LLMs often generated responses in a linear, token-by-token fashion, applying a relatively uniform computational effort regardless of the complexity of the query. This approach was akin to a human answering every question, from "What is 2 + 2?" to "Plan a family reunion for fifty people," with the same level of deliberative thought. Humans, however, operate differently. Simple questions elicit instant answers, while complex tasks demand extensive contemplation, planning, and intermediate steps before a final solution is presented.
Modern, advanced reasoning models mimic this human-like cognitive process. They are designed to spend significant computational resources on "intermediate steps" when confronted with challenging questions. Techniques like Chain-of-Thought (CoT) prompting or Tree-of-Thought reasoning enable these models to break down a complex problem into smaller, manageable parts, work through each segment, and then synthesize these intermediate findings into a comprehensive final answer. This internal process generates a substantial volume of "intermediate reasoning tokens"—a digital scratchpad of sorts. These tokens are not directly part of the "Input" (the initial query) nor are they typically presented as part of the "Output" (the final answer). This is where the legal quagmire begins: is this intermediate reasoning data actually "capital-O Output" in the legal sense? The current agreements conspicuously omit any explicit clarification.
Who Owns What: The Consultancy Analogy
The precise definitions within these agreements are not mere semantic niceties; they are the bedrock upon which ownership rights are established. OpenAI’s agreement, for instance, explicitly states that the customer "retains all ownership rights in Input" and "owns all Output," with OpenAI assigning its interest in Output to the customer. This seems comprehensive, covering both sides of the interaction. However, the lacuna surrounding intermediate reasoning tokens creates a critical vulnerability.
To illustrate this, imagine engaging a highly skilled human consultant. You hand them your company’s confidential financial forecast, asking them to derive the headcount budget for the next fiscal year. The consultant, an expert in financial modeling, meticulously works through the problem, filling an entire notebook with intricate calculations, assumptions, proprietary algorithms, and newly derived insights based on your confidential data. After hours of intensive work, they present you with a concise, one-sentence answer—the final headcount budget. They then hand you this single sentence and retain their notebook, filled with all the valuable intermediate calculations and insights. Your contract with this consultant explicitly covers the "question" (your input) and the "answer" (their output). Crucially, it says absolutely nothing about the ownership of the notebook.
This analogy precisely mirrors the situation with current reasoning models. Every time an enterprise uses such a model, it generates these intermediate reasoning tokens—a dynamic, digital scratchpad containing facts extrapolated from your prompt, internal conclusions drawn, and potentially groundbreaking insights about your business that were not explicitly stated in your input but were derived by the model’s sophisticated processing. The AI labs are acutely aware of the immense value embedded within this "notebook." OpenAI itself has publicly stated that it conceals raw chains of thought from users for a variety of reasons, including "safety, user experience, and competitive advantage." Furthermore, Anthropic, a direct competitor, bills its customers for the full extent of this internal thinking, even when none of it is ever rendered visible to them.
Therefore, enterprises are not only paying for the computational resources expended to produce this invaluable data but are also denied access to it. Moreover, the critical question of legal ownership remains deliberately unanswered. OpenAI’s "no-training" promise applies only to "Input" and "Output." The "notebook" of intermediate reasoning fits neither of these definitions. If reasoning tokens are indeed "Output," the agreements must be amended to explicitly state this and extend the same protections. If they are not, then OpenAI, and by extension other providers, have effectively created a third category of data that their agreements neither define nor protect, leaving it in a legally ambiguous and potentially exploitable state. Anthropic grapples with the identical issue; its API bills for reasoning tokens that remain unseen in the visible response, and to add another layer of opacity, it returns the raw reasoning encrypted, accessible only to Anthropic. This is not a mere drafting oversight; it is a calculated and convenient ambiguity, given the strategic and financial worth of this hidden data.
Distillation Proves the Notebook’s Immense Value
The notion that reasoning tokens are mere computational exhaust or irrelevant byproducts is easily debunked by the very conduct of the AI labs themselves. Their actions unequivocally demonstrate the profound strategic value they attribute to this intermediate data. A prime example comes from Anthropic, which highlighted an incident involving companies like DeepSeek, Moonshot, and MiniMax. These entities reportedly generated more than 16 million Claude exchanges, not for direct user interaction, but specifically to aid in training competing models. Anthropic accurately characterized this as "industrial-scale distillation"—a highly efficient shortcut to achieving advanced capabilities that would otherwise demand colossal investments in time, data, and computational power.
This practice underscores a critical point: if model outputs transfer intelligence, then the raw reasoning tokens—the detailed record of how a model arrives at an answer—represent an even richer, more granular repository of intelligence. The labs themselves confirm this. When OpenAI launched its advanced "o1" model, it explicitly stated its decision not to expose "the raw chains of thought to users" after "weighing multiple factors including… competitive advantage." This is a direct admission: the internal reasoning processes are a strategic asset, so valuable that they must be guarded from users. The labs, therefore, want it both ways: they bill customers for the computational effort of producing this reasoning, they conceal it because it is strategically valuable, and they deliberately refuse to clarify whether it legally constitutes the customer’s "Output."
The Fair Use Playbook: A Double Standard in AI
The glaring double standard inherent in this situation is impossible to ignore. The foundational premise upon which much of the generative AI industry was built rests on the argument of "fair use." This legal doctrine permits the ingestion of vast quantities of existing, often protected, human-created content—books, articles, images, code—transforming it through training, and then claiming ownership of the resulting model and its generative capabilities. The argument is that the original works are not copied verbatim but are transformed into a new asset.
Now, apply this very same logic to enterprise data. An enterprise provides confidential, proprietary material as "Input" to an AI model. The model then transforms this input into highly valuable "reasoning tokens." These tokens are not identical copies of the original Input, but they are newly generated material, derived directly from and deeply reflective of the enterprise’s confidential data. Given the industry’s historical reliance on the "transformation" argument for its own training data, why would anyone reasonably expect these same AI labs to resolve this inherent gray area against their own financial and strategic interests?
As Alex Karp aptly summarized, "the jig is up." If enterprises are not vigilant, they risk losing ownership of their intellectual alpha, the very core of their competitive edge. The individual widely perceived as erratic was, in fact, delivering a vital warning. Even seemingly robust solutions like "Zero Data Retention" (ZDR) offerings from both OpenAI and Anthropic do not entirely close this gap. While ZDR aims to ensure that customer data is not stored post-processing, it often requires a separate, explicit approval process. Crucially, ZDR typically focuses on Input and Output and does not guarantee that intermediate reasoning data is discarded rather than retained or ambiguously handled. The very fact that "not keeping your data" requires a special request strongly implies that retention, in some form, is the default operational mode—and there’s a reason for that default.
Which Is It? A Demand for Clarity
For enterprises to truly harness the power of AI without jeopardizing their core assets, a fundamental shift in contractual clarity is required. Intermediate reasoning tokens, which represent the very process of intelligent derivation from proprietary data, must be explicitly assigned to the customer. Just as one would demand the return of a consultant’s notebook containing vital calculations, an enterprise needs access to and clear ownership of these reasoning tokens. This data is not merely a record; it is a potential foundation for further reasoning, analysis, and strategic development. The AI labs have conspicuously avoided assigning these rights, primarily because doing so would necessitate relinquishing a significant source of value and competitive advantage. And if these reasoning tokens are not categorized as "Customer Content" under existing agreements, the unsettling possibility remains that they are indeed being subtly utilized to train and improve the next generation of AI models—a prospect that current lab conduct does little to assuage.
The solution is remarkably simple. Sam Altman of OpenAI and Dario Amodei of Anthropic possess the authority to resolve this ambiguity with a single, unequivocal sentence each, clearly defining the ownership and handling of intermediate reasoning tokens. Until such explicit clarification is provided, Karp’s stark warning stands: You own the prompt. You own the answer. But the most valuable part, the "notebook" of your AI’s insights and derivations from your data, they keep.
Adam Fish is the CEO and co-founder of Ditto, an edge data platform built for unstoppable operations. Its peer-to-peer sync technology keeps applications running with no cloud dependency, in U.S. defense programs including Special Operations Command and for commercial brands like Chick-fil-A.
The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

