Qwen3.8 27B scores 52 on Artificial Analysis
Points and comments are a snapshot, not live.
Qwen3.8 27B scores 52 on Artificial Analysis, matching DeepSeek V4 Flash 0731.
The model achieves an Artificial Analysis Intelligence Index score of 52, equal to DeepSeek V4 Flash 0731 (284B total parameters, 13B active). It ranks second among Qwen models, below Qwen 3.8 Max. The index aggregates nine evaluations measuring reasoning, knowledge, and agentic tool use. Qwen3.8 27B is a dense model with 27 billion parameters. Its AA-Omniscience score (knowledge reliability and hallucination) is slightly worse than Qwen 3.6, while it produces nearly twice as many output tokens per task. The model's context window is 192K tokens.
What commenters are saying
Commenters are impressed by the score relative to model size, noting it beats all medium models (40B-150B) and ties a large model ten times its total parameters. Several point out the model is token-hungry, using about 2.3x more tokens per task than GPT Luna Max, which hurts local deployment speed. A split emerges: some praise the agentic capabilities for local use, while others suspect bench-maxxing due to high variance across individual benchmarks. Skeptics recommend testing on unconventional tasks outside evals.