GLM-5.3-Flash
Points and comments are a snapshot, not live.
ZAI launches GLM-5.3-Flash, previously tested as ox-alpha, on Chinese AI chips.
ZAI released GLM-5.3-Flash, a language model previously tested anonymously as ox-alpha on OpenCode and OpenRouter. The model was served entirely on Chinese AI chips. ZAI claims a 3x improvement in end-to-end serving performance over their baseline on the same hardware, achieving hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. Pricing is $0.15 per million input tokens, $0.50 per million output, and $0.03 for cached input.
What commenters are saying
The thread highlights that serving GLM-5.3-Flash on Chinese chips is a landmark for domestic AI hardware. Commenters note the model was often slow and high-latency when free, likely due to overload. Some argue export controls have accelerated China's self-sufficiency in chips. Others suggest the efficiency claim likely compares to consumer RTX GPUs, not datacenter A100/H100. Pricing is competitive with DeepSeek V4 Flash.