ValsAI made SRE benchmark less than a month ago.
The benchmark measures can a model reverse engineer software from binaries
Yesterday GPT saturated it.
Post
Conversation
This makes me reconsider Elon's comment that AI's will code direct binary data... Still seems silly too me but at this point why not!
damn they really said "ok we make it, now beat it" and gpt just did lol
Translated from Japanese
From the graph, what you want to move comes first. Staring at performance comparisons while that's still vague—I’ve done that a few times myself and wasted so much time.
This is to me the most amazing result. I can't wait to do some reverse engineering hehehehe
The smartest way to watch college and pro football. Sling lets you do that.
The media could not be played.
less than a month to saturation is the tell for me. that timeline usually means the task distribution overlaps with training data. were the binaries sourced from public CTF repos or anything GPT might've seen before?
Discover more
Sourced from across X
Me: show me a pelican riding a bicycle
Astra:
The media could not be played.
No King rules forever.
Quote
Arena.ai
@arena
Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)!
It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing.
GPT-6 Astra by @OpenAI takes the top spot in Code Arena: WebDev with a x.com/OpenAI/status/…
Rate proposed Community Notes
From GPT-4 to GPT-6.
Sparks was a remarkably prescient paper that got a lot of pushback at the time, but absolutely sensed where the vibes were heading with LLMs based on a lot of qualitative experiments. It deserves credit in retrospect.
Quote
Adam.GPT
@TheRealAdamG
microsoft.com/en-us/research
“Sparks of Artificial General Intelligence: Early experiments with GPT-4” - March 2023