Based on my estimates, the 10,000-agent swarm that solved the Navier-Stokes problem was about 8 months of frontier progress ahead of a single GPT-6-Astra instance (with a range of 5.5 to 12.4 months)
It used 48,000 times more output tokens than the largest single model evaluation and roughly 68 times more than an entire long-horizon coding benchmark
Conversation
don't get confused that it says July here and not May
it is 8 months ahead of Astra, but Astra is 2 months ahead of trend
this roughly contradicts noam brown’s statements on the last dwarkesh pod. he sounded like he was more in the “it’s just a good model” camp and that the swarm’s value was an open question
it is a good model and with the previous model this probably wouldn't have been possible, or only with like a million parallel agents
but this breakthrough was mostly enabled by the swarm
and it will take a few generations (over Astra) until a single model can solve problems of
So the labs are sandbagging more and more to “beat” the other lab internally. Great for the public…
By your own numbers, 48,000 times the output tokens bought 8 months of progress, and 68 times the spend of a long-horizon coding benchmark. That trade is the real finding. The swarm is a way to spend compute in parallel, and the price is printed right in the post.
your 1.9 beci per 10x is fit on pro models: parallel samples plus a vote, on questions with known answers. voting saturates fast. a swarm on an open proof is buying coverage of a search space, a different curve. i'd expect the pro fit to undersell navier-stokes, not bound it
100$ in 2 years the realize that the proof is broken and using fake physics and just massive degrees of freedom to fit a math solution. Nobody will report the story, nobody will remember, the solution will stay unsolved. About 23 people can actually read what they said
10,000 agents effectively pulling ~8 months of frontier progress into the present is insane. Makes you wonder how much capability is already here, just waiting to be unlocked with enough inference-time compute.
I keep telling people in real life, you arent getting the results you want because you arent using enough inference and you arent having enough agents doing simultaneous work on the same thing. If you want to do hard work with real dollar value, a good model isnt enough, you need
Very interesting analysis !
If this prove to be accurate, it means that by end of next year, it'll be possible to run such model on local high end hardware (around 6 month lag for a 1T well scaled open source model to catch back Astra level)
We’re basically learning that enough compute can buy you a piece of the future early. The question is how fast that premium collapses.