NanoGPT Speedrun Frontier

121 points · 30 comments on HN · read original →

Points and comments are a snapshot, not live.

Prime Intellect ran 153 autonomous runs of 18 frontier models on a nanoGPT optimization benchmark.

The benchmark tasks models with improving a nanoGPT training recipe to minimize validation loss. Fable 5 achieved the best record (2,726) with an 81.7% gap closed, while Opus 5 scored 2,920 (53.6% closed). The page includes leaderboards, equal-budget comparisons, and full agent traces for 41 model-harness-seed combinations. Many models discovered similar winning ideas; differences came from how well they preserved weak signals. Some models are still running.

What commenters are saying

Several commenters questioned the benchmark's methodology and clarity. One noted the article assumes familiarity with Anthropic's evaluation setup without explaining the nanoGPT speedrun task itself. Another pointed to METR's related work on expenditure horizon and contamination concerns. A commenter highlighted that Grok's poor showing may reflect harness quality rather than model capability. Others praised GPT-5.6 Luna's cost-effectiveness for verifiable tasks.