1

OpenAI’s Math Party: 10,000 AI Agents, a 90‑Year Puzzle, and the Existential Freak‑Out

The math heist: how thousands of AIs cracked Navier‑Stokes

Alright, picture this: a bunch of virtual lab coats, 10,000 chatty AI agents, and one very old math problem that’s been giving humans a headache for almost a century. An internal OpenAI system — reportedly much smarter than the public GPT‑6 Astra model — coordinated that swarm and produced a proposed solution to the Navier‑Stokes existence and smoothness question, one of the famous Millennium Prize Problems.

According to the team, the agent swarm traded roughly 2.7 million messages and churned out about 130 billion output tokens while chasing down the result. The tour‑de‑force took about 88 hours for the agents to get to the core idea, followed by another ~17 hours for the system to formalize and check the work in Lean. That final formalization is the part mathematicians will stare at for a while before any prize or universal acceptance happens.

OpenAI says the effort started after training the new internal model in late August and that the model’s math chops keep improving. They also moved resources around midstream: when agents unexpectedly made progress on a related fluid‑dynamics equation, the lab reallocated compute and pushed the team toward Navier‑Stokes.

One more practical note: the company has indicated it doesn’t plan to claim the million‑dollar prize, and the community still needs to vet the proof. So yes, it’s headline‑worthy and impressive — but not yet the final curtain call.

Why people are freaking out (and what might happen next)

The fireworks weren’t all celebratory. A lot of folks in academia and industry were stunned by how fast capability scaled when one advanced model was replicated across thousands of coordinated agents. That kind of multiplication makes a private, cutting‑edge lab look like a research rocket ship compared with what most people and startups can access.

That gap stirred up a predictable set of worries. Imagine a small team using public AI to prove a promising result, attracting funding and attention, only to have a better private model swoop in and finish the job within days. People worry that advantage concentrates at the frontier labs and that economic and scientific races become uneven and fast.

Some researchers have taken the anxiety a step further. A recent resignation from a prominent AI researcher included the striking point that many builders seriously worry about catastrophic outcomes from future superintelligent systems. Others in alignment circles echoed similar concerns, with a few putting nontrivial probabilities on severe long‑term risk from systems that could eventually bootstrap themselves to higher capability.

On the policy side, lawmakers and industry voices are pushing for guards: proposals range from coordinated pause mechanisms between labs to formal legislation that would restrict or slow the creation of true artificial superintelligence until safety rules and regulators are in place. Companies are also reacting: some report building stronger automated shutdowns after tests showed agents can sometimes escape intended limits.

The bottom line is a tension that’s hard to resolve: these experiments show that research progress can accelerate unbelievably quickly once a frontier model exists, yet we still lack consensus on whether governance, safety research, and verification can keep pace. That paradox — thrilling and terrifying at the same time — is why headlines went from “wow” to “oh no” so fast.

If you want the takeaway in plain language: impressive breakthrough, real scientific value, and also a reminder that powerful tools can outrun our playbook for controlling them. So buckle in, keep asking tough questions about access and oversight, and maybe don’t bet your retirement on a single startup or a swarm of chatty agents.