Skip to content
Go back

Navier-Stokes: Neither Human nor AI, but the System

Published:  at  02:00 AM
Available Languages:

Navier-Stokes: Neither Human nor AI, but the System

Behind the 10,000 agents was an organization. And a human role that became scarcer without disappearing.

I believed in the plateau

I was one of those who did. Not as a matter of posture, but from reading the curves: high-quality training data appeared to be running out, gains per order of magnitude of compute were visibly flattening, and above all, chain-of-thought reasoning collapsed after more than a few steps. Models could produce a plausible answer to an isolated problem; they could not hold a line of reasoning across three hundred steps without going off the rails. I thought that wall was structural. That we would spend the next two years polishing interfaces around a capability that would itself remain stagnant.

I was wrong.

It needs to be said plainly, without immediately looking for a graceful way out. Progress continued, but not in the form I expected. Models improved. More importantly, architectures appeared that could distribute long reasoning processes, consolidate what held up, and start again elsewhere when one trajectory failed. My wall probably did exist. It was not broken down, it was bypassed, and that is where I was wrong: I had not seen that as an option.

Yet the progress of the models is not what struck me most this year. It is something else, and I use the word deliberately: what I find staggering is the way humans and machines have begun to organize together into systems of a complexity that no one truly anticipated. Not stronger models used in the same way. A different kind of apparatus.

On September 8, 2026, OpenAI published a 166-page document accompanied by a Lean formalization, claiming that an internal system had established that a three-dimensional incompressible fluid, initially at rest and subjected to a smooth force, can develop a singularity in finite time. This result has been submitted to the scientific community and still needs to be validated.

I am going to discuss the apparatus that produced it, not the result itself.

What ran for 88 hours

The story writes itself: an AI solved a problem that had remained open for nearly ninety years. An unbeatable line, and one whose very structure is false. No single model sat down in front of the problem.

The effort began on September 1, 2026, across several problems at once. The agents produced a total of roughly 4.9 million messages and 300 billion tokens, including 2.7 million messages and 130 billion tokens for Navier-Stokes alone. In parallel, close to one hundred agents spent around fifty hours working on the regularity of the unforced Euler equations, in other words Navier-Stokes without its viscosity term. Their result altered the course of the experiment: the advance was shared with the relevant groups, and more resources were concentrated on Navier-Stokes. That group grew to roughly ten thousand concurrent agents and reached its result on September 5, about 88 hours after the first launch. The groups could communicate with one another and were cross-checked through Codex. The model powering them has not been made public; OpenAI describes it as more capable than GPT-6 Astra.

Seventeen additional hours followed, during which GPT-6 Astra formalized and verified the proof in Lean. This happened afterward, on an object that had already been produced.

The numbers are spectacular, but they do a poor job of describing what happened. The agents were not simply released in parallel on the same question in the hope that one would find the answer. There were competing paths, a mechanism for circulating information between them, a selection rule, and a mid-run redirection triggered by an intermediate result.

These are not model characteristics. They are organizational characteristics.

The relevant unit is no longer the model. It is the assembly: model, agents, tools, shared memory, result circulation, selection rules, budget, and validation mechanisms. For lack of a better word, we call this a harness. It is a poor term, suggesting an accessory when this is in fact the main apparatus. What was built looks more like a miniature, synthetic scientific institution.

Notice, too, what the published numbers measure. Tokens, agents, hours, messages. Everything mechanical is counted precisely. Everything else is absent.

What humans had to decide

I am mixing two levels here, and I would rather say so explicitly. OpenAI mentions some of these decisions, but says nothing about the process behind them. The others are architectural necessities that I am reconstructing. I am not describing what happened inside OpenAI’s offices. I am describing what had to be decided for an experiment of this form to exist.

Someone had to choose which problem justified this expense. The statement needed to be important enough to warrant the effort, formalizable enough to be verifiable, and structured enough to be decomposed into parallel formulations. That three-part criterion eliminates the overwhelming majority of open questions.

Someone had to decide how many populations to run, on which formulations, and with how much diversity between them. Too much diversity and the budget scatters across dead ends; too little and ten thousand agents pay ten thousand times for the same flawed reasoning.

Someone had to define when a path dies. This may be the most delicate decision in the entire apparatus: an overly aggressive selection rule kills minority hypotheses, precisely the ones that sometimes end up paying off.

Someone had to determine what circulates between groups and in what form, a constant tradeoff between contamination and coordination.

Finally, something had to recognize that the Euler result was a signal changing the odds of success on Navier-Stokes, and reallocate resources accordingly. The public account attributes this decision to the team without saying anything about its mechanism. We do not know whether it was made by a person looking at a dashboard, by a rule written in advance, or by some combination of the two. Yet this is the most interesting moment in the entire experiment, the point where its trajectory shifts.

And there had to be mistakes beforehand. An apparatus this complex is not right on the first try. How many architectures were tested, how many budgets burned without a result, how many selection rules rewritten? Nobody knows.

This is where the most interesting imbalance in the whole affair lies. We know the number of tokens to within tens of billions. We know neither the size of the human team, nor how long they worked, nor the number of attempts. We know how to measure the compute injected into the result. We do not know how to measure the human intelligence crystallized in its architecture.

I thought this would be no more than a methodological caveat. It has since taken on a different status: among the criticisms publicly directed at OpenAI around this announcement is the claim that it underreported human involvement in the result. The company disputes the allegation, and I will not adjudicate it here. But it shows that this measurement asymmetry is not a minor editorial detail. Because only the machine component is quantified, only the machine component becomes visible, and the story of “the AI that solved Navier-Stokes” takes hold without anyone needing to manufacture it.

Degree, position, nature

What exactly changed?

AI changed in degree. Models are better, tokens are cheaper, and agents are reliable enough to chain together hundreds of steps without dissolving. That is the condition that makes everything else possible. As long as agents drifted too quickly, assembling ten thousand of them would have produced nothing. But this is the crossing of a threshold, not a change in nature: nothing that ran during those 88 hours did, at scale, something models could not already do on a small scale.

Humans did not change at all. No new capability, no newly acquired skill. They changed position. The same judgment, the same experience, the same intuition about what is worth trying, applied one level higher.

The system changed in nature. And it is the only one of the three that did.

That is what I find remarkable. None of the components mutated. Quantitative progress on one side, a simple displacement on the other, and in the end, a form that did not exist before. Neither human nor AI: the system.

In practical terms, the shift fits into one sentence. Humans set the frame and the validation conditions; machines infer, create, and execute the tasks inside that frame.

In the kind of collaboration we have all practiced for the past three years, a human formulates a task, receives an answer, verifies it, and formulates the next one. The human is in the loop on every turn, and their productive capacity is bounded by their ability to supervise. It is a hard limit: you cannot review faster than you can review.

What happened here is not the disappearance of the human from the loop. It is the human’s exit from every loop. Humans still intervene, but rarely, at moments that matter greatly: when designing the apparatus, when shifting resources, when deciding that the Euler result changes the equation. Human judgment is not removed. It becomes scarcer, moves elsewhere, and gains considerable leverage.

This is less spectacular than saying “AI does research on its own.” It is also much more demanding, because it means that the quality of the result depends on a small number of human decisions that are difficult to observe.

The question, then, is what replaced the supervision humans no longer provide. Two very different things: the apparatus’s ability to sort its own paths while it works, and its ability to produce a result that can be verified without reading it. Sort during, audit afterward.

That is the question to ask in your own field, even before asking whether the models are good enough. Without automatable selection, massive exploration does not produce ten thousand paths; it produces ten thousand things to read, and the bottleneck returns unchanged. Without an automatable audit, you very quickly manufacture an object that nobody can validate.

The interesting frontier, then, is not between human and artificial intelligence. It lies between the problems for which we know how to construct a framework for exploration and validation, and those for which we still do not know how to say what success would mean.

The real work begins now

For two years, marketing has promised AI systems that do more than execute isolated tasks. Demonstrations have so far fallen far short. This one finally gives the promise a concrete enough form to examine.

What remains is to measure what has actually been demonstrated. In a matter of days, a hybrid cognitive organization produced a mathematical object on a scale that no human team would have created in this way. That is not nothing. It may even already be a method rather than a one-off, but one whose reproducibility, true cost, and operating conditions remain unknown to everyone who was not in the room.

Nobody yet knows which exploration rules work and which ones kill diversity. Nobody knows how to preserve a minority hypothesis except by deciding to do so in advance. Nobody knows what becomes of the structure in a field where neither selection nor auditing can be mechanized.

We know that such an assembly can produce an extraordinary result. We do not yet know how much of that assembly is reproducible, generalizable, or economically sustainable.


One final word. Congratulations to OpenAI for continuing to deliver models that progress precisely where many of us, myself included, predicted a plateau.

And congratulations, above all, to the team that designed this apparatus. We do not know how many people were involved, how much time they devoted to it, or how many architectures they had to abandon before reaching this one. Their invisibility does not diminish their role. It is exactly why that role needs to be named.

Read More

Frequently Asked Questions

Did an AI solve the Navier-Stokes problem on its own?

No. The announced result was produced by a system combining models, thousands of agents, tools, shared memory, selection rules, a compute budget, and validation mechanisms, with decisive human choices about the architecture and allocation of resources.

What role did humans play in this experiment?

Humans had to choose the problem, design the agent populations, define the rules for selection and information sharing, and decide when to redirect resources. Human judgment did not disappear: it moved to a handful of rare decisions with considerable leverage.

Why is automated verification essential?

Without an automatable audit, massive exploration produces an object that nobody can validate at a comparable speed. Formalization in Lean makes it possible to mechanically check the proof's consistency after it has been produced.

What distinguishes this system from a more powerful model?

Its central properties are organizational: multiple competing paths, results circulating between groups, selection among approaches, and resource reallocation. The relevant unit is therefore the complete assembly, not the model in isolation.



Next Post
How a Mixture of Experts Manages Its Experts