OpenAI's GPT-6 Astra on ARC-AGI-3

218 points · 134 comments on HN · read original →

Points and comments are a snapshot, not live.

GPT-6 Astra scores 99.9% on ARC-AGI-3 with a specialized harness.

OpenAI's GPT-6 Astra achieves state-of-the-art scores on ARC-AGI-3, scoring 99.9% for $19K with the Provider Adapter harness and 62.7% for $26K with the Standard harness. It surpasses the human baseline in action efficiency, using fewer actions than the median human on 96% of levels. Astra develops compact symbolic world models and custom notation to track game state. The ARC-AGI-3 benchmark tests exploration, modeling, goal-setting, and planning without explicit instructions. The authors note this is meaningful progress but not proof of AGI.

What commenters are saying

Commenters debate whether these results imply AGI, with some saying goalposts will inevitably shift. Others note the counterintuitive cost curve where higher reasoning effort costs less due to fewer actions. Several point to the harness difference: the Standard harness discards context, while the Provider Adapter retains it, making the 99.9% score dependent on infrastructure. Skeptics question whether solving abstract puzzle games measures general intelligence. One commenter argues AGI is ill-defined and progress will always be met with new benchmarks.