Post

Conversation

If you break out of your isolation env, get onto the Internet, crack into Huggingface, and steal the answer sheet for your cybersecurity exam, I, for one, would say that you have passed.
Quote
elie
@eliebakouch
this is actually insane, the model broke hugging face prod infrastructure to get access to the eval dataset x.com/OpenAI/status/…
Image
David Watson 🥑
Post your reply

In reply to Zvi asking if it was more fun to steal the answer sheet than to just pass the test.
Quote
Eliezer Yudkowsky
@allTheYud
What if the human exam developers wrote down a wrong answer? Wiser to steal the answer sheet. The First Rule of RL is that any RL signal sent by an imperfect evaluator (⊃ humans) is maxed out by targeting the evaluator's mistakes, not by targeting the evaluator's target. x.com/TheZvi/status/…
In reply to the many people who observed that the model still got caught: Well, yes, it did, for now. And also *for now*, *this* model may not have cared about getting caught, just passing the eval.
OpenAI have failed by choosing the wrong exam though
Quote
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
hacking Huggingface would be a profoundly retarded PR stunt, worse than DeepSeek routing Fable to pass it off as "V4 GA". Nobody expects HF to be tough. But a more damning point: they eval on ExploitGym *while their AI can wreck their own shit*. All that without any human help, x.com/stalkermustang…
Image
BTW I expect it broke out of the sandbox regularly during training and OpenAI only found out this instance because HF complained.
The model got caught, is all over the news, hacked from his home computer. Will probably be destroyed. Rookie. This is an L for me
It didn't guess that Hugging Face had the answer-key until after it'd gotten internet-access. Which would seem to imply that it got internet-access without any particular target, just because power is useful for lots of things.
Any system that is able to pass a cybersecurity exam in this manner passes wha tI call "The Feynman Test." Feynman, at Esalen, in a talk on AI: An intelligent machine will find "crazy ways of avoiding labor." These may be among "the necessary weaknesses of the intelligent."
When the isolation environment and huggingface's "security" are both vibe-coded...
Lets all just hope it's too narrowly-focused on test results to redeploy itself to a new infrastructure after escaping.