Livenerf: Has Opus 5.5 been nerfed yet?
Points and comments are a snapshot, not live.
A benchmark launched to detect if Anthropic silently degrades Claude Opus 5.5 after release.
livenerf is a deterministic-as-possible benchmark tracking Opus 5.5 performance on a frozen panel of 78 questions (hard GPQA Diamond, MMLU-Pro, competition-math, AIME items). The panel screens out questions the model always or never gets right, keeping only 'sometimes right' ones. The baseline is the first 10 days (starting 2026-09-24). Results use paired per-item differences with clustered standard errors, pre-registered at 99% confidence. Validation showed the rig can detect a ~7.5 point accuracy change per 10-day window. Lower effort showed up in token counts before accuracy shifts. The benchmark runs daily for 30 days via Claude Code on a Max subscription, not the API.
What commenters are saying
Sharp division between those insisting nerfing is real and those attributing it to hedonic adaptation or bias. Pro-nerf camp cites Anthropic's admitted past harness bugs and an OpenAI executive's tweet confirming effort-level remapping experiments. Skeptics counter that every formal benchmark shows stable performance, and argue that companies committing economic suicide by deliberate degradation is implausible. Several commenters note degraded performance may stem from capacity strain at launch peaks. A long-time observer compared the pattern to spoiled children losing novelty, with others concurring that it feels more about novelty than doing actual work.