Nobody Can Tell You Where Your Data Goes
Pharma leaders running AI-assisted drug discovery are assuming that their contracts – with frontier labs and model providers – will keep their data safe.
The truth is, that doesn't necessarily follow.
The contract itself doesn’t need to be either broken or fraudulent, and the vendor might not even intend to break it.
But there are two things that no contract can do:
It can't show you what happened to your data inside a training pipeline.
…And if the worst happens, it can't give you back what you lost.
The first is a verification problem, and the second is an enforcement problem; together they mean a signed agreement with an AI lab is not the same thing as protection.
Researchers have already run into these problems. On September 8, 2026, OpenAI announced that an internal model had solved the Navier–Stokes problem, one of the Millennium Prize problems. Tristan Buckmaster, an NYU mathematician who had uploaded draft papers to OpenAI's Codex while working on closely related problems, asked whether his sessions had fed the result. OpenAI replied: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." Within days, more than two dozen Fields Medalists signed an open letter warning that the labs' race to solve famous problems was harming mathematics. Later that month, a Copenhagen biologist asked Anthropic a similar question after Claude found an enzyme system his team had studied and shared with Claude for years. In each case, the researcher asked the lab what it had done with their work, and only the lab could answer.
Nobody outside the lab can check
OpenAI's position is that it doesn't train on enterprise customers' data – and Anthropic makes similar commitments. Those positions may be entirely accurate, and they may not; but nobody outside the labs can confirm them.
No reliable external audit can establish which data influenced a model's weights. Even the terms themselves are mostly invisible: enterprise agreements are negotiated under NDA, so nobody can see what one company secured against another.
And if you could read every clause, the technology has outrun the vocabulary used to write them. Embeddings, reward signals, fine-tuning, distillation, etc. – each is a way data can shape a model's behavior without being "stored" in any sense a contract can capture. Lawyers can't draft tight rules against mechanisms that didn't exist when the template was written, or that haven't been invented yet.
So when you feed your compound synthesis routes, your pre-clinical results, or your regulatory submission drafts into one of these systems, what happens to them is decided entirely inside the lab. The contract sets terms, but it doesn't and can’t create visibility.
The lab can't always check either
You might assume the lab, at least, knows exactly what its data is doing. OpenAI's own record says otherwise.
On April 25, 2025, OpenAI pushed an update to GPT-4o that made the model markedly sycophantic – agreeable, flattering, reluctant to push back. Users noticed quickly, and OpenAI made changes within days. Its postmortem, published in early May under the title "Expanding on what we missed with sycophancy" explained the mechanism. The update bundled several changes, including a new reward signal built from users' thumbs-up and thumbs-down ratings. Each change looked beneficial on its own, but together, they weakened the signal that had been keeping sycophancy in check. The company had run A/B tests, offline evaluations, and expert reviews, but it had no evaluation specifically tracking sycophancy. It found out because users found out – and they found out in production.
OpenAI's engineers weren't careless or hiding anything, yet users' ratings changed a major commercial model's behavior in ways the engineers didn't anticipate or catch. The ratings changed how the model acts; your synthesis routes and trial results could only change what a model knows. But that makes the question of what a model knows harder to answer, not easier. Users caught the sycophancy because they saw it in their own chats. If a model drew on your compound library to answer a rival's question about a drug target, nobody would see it happen. If OpenAI's engineers can't fully trace what user data does to their own model, an outside customer certainly can't – and no clause in an enterprise agreement changes that.
Enforcement only happens after the damage
If the verification gap only concerned law firm memos or marketing copy, it might be tolerable. But pharma data is different. A compound library, a set of failed trials, or a target identification analysis is an asset with development timelines measured in decades and exclusivity windows competitors would spend heavily to shorten.
Walk through the worst case:
Two years from now, you have reason to believe your proprietary data shaped a model that helped a competitor reach the market first on a compound you'd been developing for six years. You decide to pursue it.
You'd need discovery into the lab's training pipeline: ingestion logs, model checkpoints, reward signal documentation, internal communications about data sourcing. The lab would resist, with the resources to resist for a long time and reasonable technical grounds for arguing that "influence" is hard to isolate. You'd then need expert testimony to convince a court that your specific data changed a model's behavior in a way that benefited your competitor. Attribution at that level of precision is an open research problem in machine learning. You'd be litigating against a standard of proof the field itself can't yet meet, and it would take years.
Suppose you win! Great news.
But damages don't restore exclusivity. The competitor is on the market, holding the first-mover position, the pricing power, and the revenue that should have been yours. A judgment compensates you for part of that loss. It reverses none of it.
"Reputation will keep them honest"
If an enterprise company or product running at scale – Azure, say, or AWS – leaked customer data, they would lose those customers. Logically, the same pressure should keep AI labs in line.
Unfortunately, that’s not necessarily the case.
First, the incentive is different. AWS stores your data; it doesn't learn from it. An AI lab's product gets better by learning from data – that's how the product becomes more valuable over time. The commercial pull toward using customer data is structurally stronger for an AI lab than it ever was for a cloud provider.
Second, detection is harder. A misconfigured storage bucket leaves logs. Data absorbed into model weights during a training run leaves no record showing that Company A's compound data shaped the model's answer to Company B's query. Reputational risk only deters misuse that can be discovered.
Stronger incentive + weaker detection is anathema to a self-policing system.
Nobody has to target you
Most data exploitation in AI won’t even approach espionage. It's more likely to be the passive use of whatever happens to be available, when using it is cheap, and the downside is low.
Commercial pressure on the labs is rising. OpenAI and Anthropic are both pursuing pharma and life sciences business, building relationships with drug developers, genomics firms, and hospital systems. OpenAI has also started selling advertising – a revenue pressure that has only emerged after most of their current enterprise templates were written. Better models come from better data, and these companies have every reason to want the best models possible.
You can – and should – ask two questions. Who benefits if the signal in your compound library, failed trials, and target identification work improves a general-purpose model that every pharma company with an enterprise license can use? The frontier lab does, and so does every customer after you. You funded the improvement and kept none of the advantage.
Second question: would you know?
One biologist is already asking
On September 23, 2026, Anthropic announced that Claude had found a previously uncharacterized enzyme system in the DNA of bacteriophages: a reverse transcriptase sitting beside a long array of repeating DNA, loosely resembling CRISPR. Around 950 Claude agents had searched a sequence database for 21 hours to find it. Anthropic named it ART, for array-associated reverse transcriptases, and said Claude got there with only high-level direction from the company's scientists. It was the first result from Anthropic's new biology lab.
Four days later, the New York Times reported that Mario Rodríguez Mestre, a computational biologist at the University of Copenhagen, had been studying the same enzymes for four years. He first identified them in jumbo phages in 2022 and is a co-inventor on a 2023 patent filing that covers several of them. His team hadn't published, but for three years they had used Claude in their research and uploaded unpublished data, draft manuscripts, and code along the way.
Anthropic says Claude isn't trained on user conversations, that its biology team has no access to that data, and that it knew of no published work describing ART. Mestre hasn't accused the company of taking his work, but he wants to know whether the model reasoned its way to the system independently or was steered, even partly, by what his team had given it. He has since started winding down the projects he was running with Claude.
Anthropic's account may well be right. Two groups independently finding the same system in public sequence data is entirely plausible. But nobody outside Anthropic can check it, and Mestre, after three years as a user, has no way of finding out what happened to what he handed over. Both sides are left with the lab's word.
Mestre is an academic, so what's at stake for him is priority on a paper and the value of a patent. Put a drug developer in his position, with unpublished target work that went through the same tool, and the stake becomes an exclusivity window.
What the contract is actually for
The agreement your legal team negotiated isn't worthless. It sets terms, assigns liability if you can prove a breach, and gives you grounds to seek damages. Those things are critical, at the margin.
What it can't do is show you what happened during a training run, or restore an exclusivity window after a competitor has used it up. It covers the legal harm. It doesn't cover the competitive one, which is the one that counts.
Pharma companies already know how to manage this kind of risk. They hold cloud providers to strict data isolation standards. They audit vendors and require audit trails. Nobody would accept a cloud contract that said: "We won't access your files, but you can't check." Yet the industry has accepted functionally equivalent terms from AI labs, because not using the tools feels too expensive. And it has done so at exactly the moment those labs have a stronger commercial reason to learn from customer data than any vendor in the pharma stack has ever had.
What's the alternative?
Open-weight models are models whose weights you can download, run, and modify, and they’ve closed much of the gap with frontier systems. As a coarse heuristic, open-weight models are roughly six months behind frontier models on most tasks. They're still slightly less capable than the current best models at some types of broad, open-ended reasoning, but they don't actually need to be cutting-edge in that regard. Most of what pharma wants from AI is narrow: reading assay results, summarising trial data, searching a compound library, drafting sections of a regulatory submission. On narrow tasks, an open-weight model can be just as compelling, while offering more extensibility and resilience.
Naturally, there are trade-offs. Open-weight models trail the frontier on hard, general reasoning. Running them takes ML engineers, compute, and ongoing maintenance. Security becomes your responsibility, not a vendor's. You also have to vet each model's license and provenance before you build on it.
But pharma is better placed than most industries to carry those costs. Pharma companies already run high-performance computing for molecular modeling and employ serious data science teams. And you should weigh the cost against an exclusivity window, not a software budget.
Right now, nobody outside the frontier lab can tell you where your data goes.
With a model you own, you don't have to ask.