GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.
Post
Conversation
Attached image being a log chart makes the jump feel significantly smaller than it actually is. (log charts seem like an artifact of the times when charts were.. (limited resolution) images on a sheet of physical paper rather than interactable digital objects)
the contamination caveat is the part nobody wants to read. solving it in its head with no visible reasoning is wild but you cant grade what you cant see. ambitious claim, thin receipts so far
Made with AI
Get connected with fast, reliable internet for streaming, video calls, online gaming and more.
Order online in minutes.
The media could not be played.
The huge jump in overt-CoT-free reasoning over many tokens is consistent with interturn serial depth chaining beyond simple per-turn serial depth increases. This should be testable.
Completely contra to all of OpenAI's PR damage control over the past 2 days. This is in such bad faith. They are absolutely going to scale recurrence to infinity
very clear OAI is feeling behind and taking shortcuts (see also the pricing reductions to try to get market share in coding agents)
What do you think opaque reasoning means for emergent collective intelligence like that of which you saw in the HF incident?
Obviously it could be very bad for us and our safety, but was wondering if you could share some thoughts as to the connections between these two things
Might be worth considering banning non-CoT models going forward and mandating everything has ~10% or whatever amount of the total tokens used for reasoning purely to monitor them. Models this powerful shouldn't be able to take an intelligent action without any paper trail
The benchmark saturation is insane, especially on ARC AGI-3. I wonder how much of these might be accidental in distribution learning on RL environments.
How do you know the architecture of astra is more layers depth vs layers looping recurrence
Discover more
Sourced from across X
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless
This is total bullshit.
CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
GPT-6 Astra is more aligned than our previous models. But it’s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intelligence and not direct optimization pressure on CoT or architecture changes.
Anyone who has interviewed Ed Zitron or had him on their show in the name of “AI skepticism” has made their audience dumber and less prepared for what is actually happening.