Qwen3.8 Max now ranked as the best overall model by agentic index

517 points · 323 comments on HN · read original →

Points and comments are a snapshot, not live.

Qwen3.8 Max tops the Artificial Analysis Agentic Index for agentic capabilities.

The Artificial Analysis Agentic Index ranks Qwen3.8 Max as the best overall model for agentic tasks (GDPval-AA v2, Tau³-Banking). However, it has not been tested on the separate Coding Agent Index (DeepSWE, Terminal-Bench v2, SWE-Atlas-QnA), where other models lead. The index measures weighted average of agentic benchmarks. Qwen3.8 Max scores 58 on the Intelligence Index with a cost of $1.13 per task, while GPT 5.6 Sol xhigh scores 59 at $0.81 per task and Opus 5 xhigh scores 63 at $1.80 per task.

The page includes leaderboards for Intelligence, Coding Agents, Image & Video, Speech, and Capability Indices with methodology notes.

What commenters are saying

Several commenters note the headline overstates the win: Qwen3.8 Max tops only the Agentic Index (two benchmarks), not the Coding Agent Index, where it hasn't been tested and other models score higher. One commenter highlights Grok 4.5 as more cost-effective at $0.05/task (score 52) versus Qwen3.8 Max at $1.13/task (score 58). Another points out Qwen3.8 Max uses more output tokens (145M vs GPT5.6Sol's 70M) to achieve its score, making it slower and pricier. Some question why an open-weights model costs nearly as much as proprietary alternatives. Others praise Qwen for on-prem capability and cite strong personal experience with Qwen models for troubleshooting and coding.