Benchmarking Opus 5 on SlopCodeBench

345 points · 90 comments on HN · read original →

Points and comments are a snapshot, not live.

Article body wasn't reachable. The HN discussion summary is below.

What commenters are saying

Commenters broadly agree Opus 5 is a modest step up from Opus 4.8, not revolutionary like Fable once felt. Several note it tends toward pedantic, overly verbose outputs ("slop"), causing feedback loops in agentic workflows. Some prefer Codex Sol or Opus 4.8 for daily coding. The thread explores deeper questions: how to benchmark maintainability and state-space complexity in AI-generated code, and how to design agent frameworks (CLI tools, state machines) that prevent drift. A known failure mode: hidden prompts that encourage pedantic behavior.