GPT-6 Astra on robot arms
Points and comments are a snapshot, not live.
GPT-6 Astra beats Claude Fable 5.1 in block bowl task by 95% vs 40%.
OpenAI's GPT-6 Astra controlled YAM robot arms on two tasks: placing a red block into a bowl and inserting a blue puzzle piece into a groove. On the bowl task, Astra succeeded in 19 of 20 trials (95%) at $0.94/run in 2.5 minutes, versus Fable 5.1's 8/20 (40%) at $2.12/run in 6.8 minutes. On the puzzle task, both models completed only 2/20 trials; Astra reached the groove but stalled at the final insertion step. Astra used far fewer output tokens per run (2.1k vs 12.9k for Fable 5.1 on bowl).
What commenters are saying
Commenters largely found the results impressive, praising the clear methodology and transparency. Some expressed disappointment that robotics progress remains slow on mundane tasks like folding laundry or cooking, despite LLM advances in coding. Several noted hardware reliability as the key bottleneck, not software intelligence: robots break down frequently, unlike the massive datasets available for text. A few discussed potential for LLM-powered self-driving cars, but concerns about prompt injection attacks were raised.