Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Points and comments are a snapshot, not live.
Claude Opus 5.5 Max offers same intelligence as Opus 5 at half the cost per task.
Artificial Analysis benchmarks Claude Opus 5.5 on its Intelligence Index v4.3.2, which aggregates 10 evaluations including AA-Briefcase v1.1, SciCode, Humanity's Last Exam, and AA-Omniscience. The model achieves comparable intelligence scores to Opus 5 at roughly half the weighted average cost per Intelligence Index task. The page tracks performance across reasoning levels (medium, high, max) and compares cost efficiency, token usage, and context window limits. It shows pricing per million tokens for cached, input, and output tokens, and breaks down cost by token type per evaluation task.
What commenters are saying
The dominant reaction highlights the cost improvement: half the cost per task compared to Opus 5 at high effort. Some users question whether metrics are meaningful without weighting by task success rate. A top comment warns that Max reasoning can exhaust the 128,000-token budget on simple requests like generating an SVG of a pelican, never reaching output. Users report inconsistent performance across providers over time, with one noting Sol's bug-finding rate dropped 50% in two weeks. Responses suggest stochastic variation, API changes, or user habituation as possible explanations.