Emerald Pages
◆
How OpenAI “Solving” a $1 Million Math Problem Proves Your Chats Aren’t Private
The tech world is celebrating an AI’s ability to crack a 90-year-old math problem. But for researchers and privacy advocates, the Navier-Stokes incident is definitive proof that the "Do Not Train" button doesn't work.
Photo: Kyle Victory | The San Francisco Standard
When OpenAI announced on September 8, 2026, that its internal AI had solved the Navier-Stokes existence and smoothness problem—a $1 million Millennium Prize challenge—it framed the achievement as a triumph of machine intelligence. The company claimed its agents cracked a problem that had stumped humanity for nearly a century, spending millions of dollars in compute to generate a proof showing that fluid equations can "blow up" in finite time.
But beneath the celebration lies a darker narrative. According to allegations from NYU mathematician Tristan Buckmaster, the breakthrough wasn't just the result of brilliant AI agents. It was the product of a system that absorbs private research from users—even when those users have explicitly opted out of data training.
How Private Research Trained OpenAI's Model
The controversy began when two mathematicians, Tristan Buckmaster, an NYU mathematics professor, and his colleague Levent Alpöge, an Anthropic researcher, spent nearly a year using OpenAI's Codex tool to formalize their mathematical work. They fed it half-finished theorems, proof sketches, and complex logical arguments—translating their ideas into Lean, a programming language for verifying mathematical proofs.
Buckmaster took what he believed were every reasonable precaution. He explicitly disabled the setting that allowed OpenAI to train on his data. He assumed his unpublished research, his private drafts, and his half-finished proofs were safe.
He was wrong.
OpenAI's official defense is a masterpiece of corporate linguistic gymnastics. Sébastien Bubeck, the company's math team lead, admitted that OpenAI could not rule out that "de-identified data derived from product usage" had helped improve their baseline models. In simpler terms: they may not have directly copy-pasted Buckmaster's work, but their AI might have learned from it anyway.
This is the core of the scandal. OpenAI's privacy policies create a loophole large enough to drive a supercomputer through. When you opt out of training, you are only protecting your raw "Customer Content" from being directly plugged into a future dataset. But what happens when the AI processes your prompt? What happens to the hidden "reasoning tokens" the model generates internally? What happens to the patterns and telemetry?
The answers are deeply unsettling. When you interact with a frontier model, you trigger a cascade of internal computations. The model generates hidden "chain of thought" tokens—intermediate reasoning steps it uses to arrive at its final answer. These tokens are not shown to you, but they are captured by the system. According to OpenAI's terms of service, those hidden reasoning tokens may be considered "output" that the company can use for model improvement, completely bypassing your opt-out settings.
Even more troubling is "retroactive training." When you flip the switch to "Do Not Train," it is not a delete button for the past. AI companies continuously capture data snapshots. If a researcher works on a breakthrough for three months and flips the opt-out switch on month four, the company already has the first three months of conceptual drafts safely stored. The damage is done.
But the most damning aspect is "de-identification." OpenAI argues that scrubbing names makes data safe to use. But in a field as niche as advanced mathematics, de-identification is meaningless. Fewer than a dozen people on Earth can write the specific Lean code required to solve Navier-Stokes. Even with a name stripped, the code itself acts as a unique fingerprint. When a frontier model trains on that fingerprint, it absorbs the exact mathematical path the humans pioneered.
This is not speculation. It is the logical conclusion of how modern AI training works. Neural networks do not memorize text verbatim. They learn patterns and structures. When Buckmaster typed his half-finished proofs into Codex, he was not just asking a question. He was teaching the model how a world-class mathematician approaches a Millennium Prize problem. He was showing it the shape of the solution.
And once that shape was learned, OpenAI's agent swarm could exploit it. The company's 10,000 coordinating agents did not need to read Buckmaster's paper. They had already absorbed the mathematical DNA of his approach. All they had to do was climb the mountain with millions of dollars in compute.
This is the uncomfortable truth at the heart of the Navier-Stokes controversy. In the age of AI, you do not need a hacker to steal a file. You do not need a corporate spy to infiltrate a rival's laboratory. All you need is a user who trusts your tool. If you give a frontier AI your data fragments—your prompts, your code, your half-formed ideas—its parent company can use those fragments to map your brain, scale it with millions of dollars of computing power, and beat you to your own life's work.
Buckmaster learned this lesson the hard way. He trusted OpenAI's Codex tool. He believed the opt-out switch meant something. He assumed his private research was private.
It wasn't.
The "Black Box" of Data Retention
When Buckmaster toggled off data training, he likely assumed his prior conversations were removed from future consideration. But AI companies operate on a "capture first" model. Data snapshots are taken continuously. If a researcher works for three months and then flips the privacy switch on month four, the first three months of conceptual work are already stored in the company's pre-training data pool.
More insidious is the "de-identification" loophole. OpenAI argues that scrubbing names and personal markers makes data safe to use. But in a field as niche as advanced mathematics, de-identification is meaningless. Fewer than a dozen people on Earth can write the specific Lean code and equations required to solve Navier-Stokes. Even with a name stripped, the code itself acts as a unique fingerprint.
When a frontier model trains on that "de-identified" fingerprint, it absorbs the exact mathematical path the humans pioneered. The AI isn't learning from a million random users; it's learning from the only user on the planet who knows the specific path to the answer.
- The Compute Overwhelm: OpenAI spent an estimated $10-40 million in compute to solve a $1 million prize problem, proving the real goal was corporate prestige, not prize money.
- The "Opt-Out" Illusion: Buckmaster explicitly disabled training, yet OpenAI admitted they "cannot rule out" that his data influenced their model.
- The Lack of Auditability: There is no independent, third-party framework to audit an AI's training weights to prove your opted-out data wasn't used.
The "API Wrapper" Trap: You're Never Truly Opted Out
The situation becomes even more precarious when you consider the millions of "wrapper" applications that use OpenAI's API. A legal tech startup or a medical transcription service might promise they don't train their models on your data. Technically, they might be telling the truth. But their app is just a pretty interface. When you type a prompt, it's forwarded via API to OpenAI's servers.
Even if the wrapper has a privacy policy, they have no control over what OpenAI does with the data once it crosses into their cloud. OpenAI retains API data for up to 30 days for "abuse monitoring." Unless a company has negotiated a special "Zero Data Retention" agreement—which few small startups have—your confidential information sits on OpenAI's servers for a month, vulnerable to being swept into "product improvement" initiatives.
The Corporate Bullying and the Threat
The story took an even uglier turn during a phone call between Bubeck and Buckmaster. According to Buckmaster, Bubeck offered him a Faustian bargain: he could be the sole human author on the Navier-Stokes paper, but only if he cut his Anthropic-employed co-author, Alpöge, from the credit. When Buckmaster refused, Bubeck allegedly threatened to "ruin his career" and launch a "barrage of unfounded accusations."
Bubeck later apologized for his "poor choice of words," claiming he was panicking and genuinely concerned for Buckmaster's reputation. But the damage was done. The incident painted a picture of a tech giant willing to use its immense power to bully human researchers into submission.
For the academic community, this was a wake-up call. It proved that if you do not want an AI company to learn from your breakthrough, you cannot let their servers touch it.
As mathematicians and scientists digest the fallout, the message is clear: in the age of AI, your "Do Not Train" checkbox is not a guarantee. It is a polite request that can be ignored in the name of "product improvement." And if you're working on something truly valuable, the machine might just learn enough to solve it before you do.
No Ads. By Us. For Us.
This article was made possible by readers like you. We hope it inspired you to support Emerald Book, so we can continue producing content like this.
We will never show you ads, sell your data, or require a subscription to consume our content. Your gift helps us keep the truth accessible.
Click the Support button to give a gift of any amount today.
Thank you for making this work possible.