GPT-6 Astra on robot arms

212 points · 170 comments on HN · read original →

Points and comments are a snapshot, not live.

GPT-6 Astra beats Claude Fable 5.1 in block bowl task by 95% vs 40%.

OpenAI's GPT-6 Astra controlled YAM robot arms on two tasks: placing a red block into a bowl and inserting a blue puzzle piece into a groove. On the bowl task, Astra succeeded in 19 of 20 trials (95%) at $0.94/run in 2.5 minutes, versus Fable 5.1's 8/20 (40%) at $2.12/run in 6.8 minutes. On the puzzle task, both models completed only 2/20 trials; Astra reached the groove but stalled at the final insertion step. Astra used far fewer output tokens per run (2.1k vs 12.9k for Fable 5.1 on bowl).

What commenters are saying

Commenters largely found the results impressive, praising the clear methodology and transparency. Some expressed disappointment that robotics progress remains slow on mundane tasks like folding laundry or cooking, despite LLM advances in coding. Several noted hardware reliability as the key bottleneck, not software intelligence: robots break down frequently, unlike the massive datasets available for text. A few discussed potential for LLM-powered self-driving cars, but concerns about prompt injection attacks were raised.