Post

Conversation

GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems. This seems extremely concerning! That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination). Related to this, UK AISI found Astra has much worse monitorability. I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.
Image
David Watson 🥑
Post your reply

Attached image being a log chart makes the jump feel significantly smaller than it actually is. (log charts seem like an artifact of the times when charts were.. (limited resolution) images on a sheet of physical paper rather than interactable digital objects)
Square profile picture
the contamination caveat is the part nobody wants to read. solving it in its head with no visible reasoning is wild but you cant grade what you cant see. ambitious claim, thin receipts so far
Made with AI
Square profile picture
Get connected with fast, reliable internet for streaming, video calls, online gaming and more. Order online in minutes.
The media could not be played.
The huge jump in overt-CoT-free reasoning over many tokens is consistent with interturn serial depth chaining beyond simple per-turn serial depth increases. This should be testable.
Completely contra to all of OpenAI's PR damage control over the past 2 days. This is in such bad faith. They are absolutely going to scale recurrence to infinity
very clear OAI is feeling behind and taking shortcuts (see also the pricing reductions to try to get market share in coding agents)
What do you think opaque reasoning means for emergent collective intelligence like that of which you saw in the HF incident? Obviously it could be very bad for us and our safety, but was wondering if you could share some thoughts as to the connections between these two things
If it's better at reasoning in general, then it will also be better at opaque reasoning. My question is how much not using the CoT degrades its capabilities.
Might be worth considering banning non-CoT models going forward and mandating everything has ~10% or whatever amount of the total tokens used for reasoning purely to monitor them. Models this powerful shouldn't be able to take an intelligent action without any paper trail
The benchmark saturation is insane, especially on ARC AGI-3. I wonder how much of these might be accidental in distribution learning on RL environments.

Discover more

Sourced from across X
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless This is total bullshit. CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
GPT-6 Astra is more aligned than our previous models. But it’s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intelligence and not direct optimization pressure on CoT or architecture changes.
Image
btw whatever this magical classifier technology is i would really like to know. if you can somehow train against a CoT monitor and avoid deception and evasion that’s a big deal, and it’s included in this post as a throwaway line!
Image
Anyone who has interviewed Ed Zitron or had him on their show in the name of “AI skepticism” has made their audience dumber and less prepared for what is actually happening.
Quote
Dan Luu
@danluu
How accurate have Ed Zitron's AI skeptic predictions been? danluu.com/zitron/
    Feb 2024: "I believe we're reaching the upper limits about what generative AI can do and how accurate its outputs can be."
        Wrong4
    March 2024: "Have We Reached Peak AI?"; another prediction that hallucinations mean that AI progress is limited to then-current levels
        Wrong
    April 2024: "As I previously warned, artificial intelligence companies are running out of data ..."; another prediction that models can't improve because there's no more data
        Wrong
    June 2024: OpenAI growth is stalling (with the implication it will continue to stall), which will lead to some kind of collapse of OpenAI
        Wrong (it could be the case that OpenAI will collapse but, if so, it won't be due to any kind of growth stall from 2024)
    July 2024: "Generative AI, as I said back in March, is peaking, if it hasn't already peaked. It cannot do much more than it is currently doing, other than doing more of it faster with some new inputs"
        Wrong
    July 2024: "Gen...