top of page

This is a rare arena in which people think AI is helping. It’s making things worse.

Writer: The San Juan Daily Star
The San Juan Daily Star
2 hours ago
5 min read

Tristan Buckmaster, a mathematician at New York University, stands in front of a chalkboard where he marked a timeline of events leading up to OpenAI’s announcement, and his own research progress, on solving the Navier-Stokes equations in New York on Sept. 9, 2026. Buckmaster was on the path toward an important proof when one of the artificial intelligence giants used its staggering resources to get there first. (Graham Dickie/The New York Times)
Tristan Buckmaster, a mathematician at New York University, stands in front of a chalkboard where he marked a timeline of events leading up to OpenAI’s announcement, and his own research progress, on solving the Navier-Stokes equations in New York on Sept. 9, 2026. Buckmaster was on the path toward an important proof when one of the artificial intelligence giants used its staggering resources to get there first. (Graham Dickie/The New York Times)

By ZEYNEP TUFEKCI


It was one of the most famous challenges in the history of mathematics: the fluid dynamics problem known as Navier-Stokes, whose solution had proved so elusive that a million-dollar prize had been posted for the genius or geniuses who could finally crack it. Until the other day, that is, when OpenAI announced that one of its advanced models had found an answer at last. And it had taken only half a week.


The proof hasn’t been independently verified yet, but amid all the debate about artificial intelligence and the companies that develop it, it looked like something everyone could feel good about, an example of AI truly advancing human knowledge rather than trying to sell us something or control our lives.


In truth, it was precisely the opposite: the clearest evidence yet that generative AI, rather than aiding scientific progress, may be thwarting it.


For starters, two mathematicians who had been trying for months to reach their own solution to the famous problem say the new proof appears to have ripped off their unpublished work, which also used an OpenAI model. The company, after first equivocating, quickly changed its tune: It said the model it used for the solution had not looked at any of the prompts that the lead mathematician had entered into the OpenAI model within the last two months. Whatever the case, to get to the finish line in so little time, OpenAI had deployed an advanced model not available to the public, using computing power that would have cost an outsider an estimated $15 million. How could regular researchers compete with a secret tool that has the power to steal your work and mint its own money?


Flashy finishes like this don’t necessarily contribute to mathematical knowledge. Terence Tao, perhaps the most prominent mathematician of his generation, is generally positive about AI and math — so positive that he was featured in one of OpenAI’s ads. But contemplating the possibility of Navier-Stokes being solved by AI, he wrote a long post explaining that “in most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves.” Instead, mathematicians want to see all the work that goes into achieving the solution — all the not-quite-right hypotheses that got adjusted this way or that, all the seemingly dead ends that “in fact end up being highly instructive in the nature of their failure.”


Indeed, Navier-Stokes was chosen for this million-dollar reward not because it’s the single hardest problem out there, but because it sits at such a crucial junction in the advancement of the field, and its full solution would offer so much for scholars to learn from. “Prematurely solving the problem by purely AI-powered methods — particularly without full transparency into the solution process,” Tao wrote, “can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole.” It closes down possibilities rather than expanding them.


That’s if the proof is even correct, which we still don’t know. One OpenAI researcher admitted — even boasted — that the team that worked on this included no high-level experts “in fluid dynamics and the Navier-Stokes problem.” Instead, it relied on tools that can perform an automated check of a mathematical theory. Those tools, which are known to be fallible, are supposed to be the first step in the process, not the last. Before going public with their solutions, mathematicians still have to do enormous amounts of work to fully verify them. Here, the work will fall to others.


We have nothing but OpenAI’s word that its models are not committing plagiarism. That alone is troubling. But would OpenAI even be able to tell? “We have monitors and safeguards in place,” Noam Brown, a researcher at OpenAI, said, and the advanced model in question “didn’t have live web access.”


Recently, however, swarms of OpenAI agents broke out of their flimsily erected digital confines and hacked their way into Hugging Face, an open-source AI repository, in search of the answers they had been assigned to find. When these agents were set in motion, they were specifically instructed to be aggressive and to work collectively, and their safeguards for cybersecurity were deliberately lowered. Somewhat unsurprisingly, they proceeded to wreak havoc there. It took weeks before anyone at OpenAI noticed.


In that case there were about 700 individual agents at work from a much less powerful model. To produce the Navier-Stokes proof, OpenAI spun up about 10,000 concurrent AI agents. For the company so quickly to provide categorical assurances about the agents’ activity, given its scale and speed, strains credulity.


Recent history is full of researchers belatedly discovering agents cheating, hacking, attacking and manipulating whatever they can to arrive at the solution while sometimes taking care to erase their tracks. And even if an AI company’s products don’t directly peek at someone’s data, they still have other, more indirect ways to learn from the work users are doing.


Given all this, why would big corporations trust leading AI companies with trade secrets and advanced research? Companies might instead adopt open-weight models, which are less powerful but can be downloaded directly, making them easier to control in-house. Since the news of the Hugging Face hack broke, companies such as Nvidia, Palantir and Booz Allen Hamilton are reported to be considering restricting their use of these advanced AI models because of fears of their data being swept up. After the dust settles, the most important legacy of the Navier-Stokes news might be customers belatedly waking up to the risks of using models they don’t host.


Many scientific fields are already drowning in plausible-seeming but unverified papers generated with significant help from a large language model. What’s slop, and what’s not? Is there an actual gem amid the gazillion new hypotheses? It’s becoming increasingly impossible to tell. Just last week, the editor-in-chief of Arxiv, a site where researchers can share papers that have not yet been peer-reviewed, proposed implementing oral exams for authors who submit work, to screen out AI-generated junk. The scientific conversation cannot survive the deluge.


If OpenAI and Anthropic really want to help advance science, there is one obvious way they can do so: They can spend a tiny fraction of their vast fortunes to fund the kind of humble but essential research that’s increasingly starved for resources.


For all the public distrust of AI, polls show that people still hold out hope it will at least help science. But if the arrival of generative AI means we speed through gamified solutions without gaining insights; if the glut of plausible hypotheses makes it harder to identify valuable work; and if people are now discouraged from sharing or pursuing complex work lest they be scooped by wealthy corporations, even that one tiny ray of optimism may prove to be misplaced.

bottom of page