I havent updated this benchmark in a while. Astra completely saturates my spatial reasoning eval. I am at a loss for words, and i'm declaring LLM vision solved. Every so often when a new model came out I would test it on a sample question and it failed. I tested Astra and it kept getting answers correct, so I evaluated it...
Quote
spicylemonade
@spicey_lemonade
Gemini 3 achieves state of the art performance in SpatialBench. A Spatial reasoning benchmark for VLM to test their tracing, 3d visualization ability, and reasoning across each.