GPT-6 Astra consistently scores better than MolmoAct2 across all five bimanual robot tasks. In 99 of its
100 trials, Astra scored at least as high as MolmoAct2’s best trial on the same task.
MolmoAct2 could lift the bowl but did not attempt pouring the paper clips. Astra poured the paper clips 30% of the time (6/20), but always spilled the paper clips.
MolmoAct2 touched the ruler but did not succeed in lifting the ruler. Astra could do mid-air transfers 10% of the time (2/20) and moved it to the other cup 5% of the time (1/20).
MolmoAct2 did not approach the notes pads, while Astra touched the pads in all 20 trials, and succeeded in removing the middle pad without the top pad touching the table 1 out of 20 times.
Our results across 200 trials show that a LLM outperforms the SOTA open-source robotics VLA on five bimanual tasks. Phillip Isola calls this the beginning of "Robot-Use Agents".
Quote
Phillip Isola
@phillip_isola
·
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots.
I wrote a short blog post with my thoughts on the advent of these "robot-use agents."
https://web.mit.edu/phillipi/www/writing/robot-use-agents.html…
I think it's an important change in the trajectory of robotics!