GLM-5.3: Frontier coding with emergent cyber capabilities

1111 points · 545 comments on HN · read original →

Points and comments are a snapshot, not live.

Scaling post-training on GLM-5.3 produces gains in coding and emergent cyber exploitation capabilities.

GLM-5.3, built on the same base as 5.2, achieves gains entirely from post-training RL scaling. It scores 28.3 on Terminal-Bench 3.0 (up from 4.6) and 66.9 on DeepSWE v1.1. On cyber benchmarks, it scores 84.5 on CyberGym and 54.4 on ExploitBench, more than doubling 5.2's 24.4. In collaboration with security teams, the model identified 2,436 vulnerabilities across 269 open-source projects, with 1,097 medium-to-high severity. The model uses the open-source slime framework for RL scaling. Weights will be released in two weeks after safety evaluation.

What commenters are saying

Commenters largely focused on the implications of open-weight models with unrestricted cyber capabilities. Several noted that Chinese labs are making frontier-level models freely available, while US vendors restrict access to their best models for security work, forcing defenders to use open alternatives. Others argued that the gap to Anthropic's Fable and OpenAI's Sol remains, but defense teams may not have access to those models anyway. Some pushed back on the notion that Anthropic can do more, noting government intervention shut down Fable for a month.