Xiaomi Mimo 2.6 live post-training dashboard

485 points · 139 comments on HN · read original →

Points and comments are a snapshot, not live.

Xiaomi live-streams Mimo 2.6 post-training RL dashboard showing two concurrent runs.

Xiaomi's Mimo team is live-streaming the post-training reinforcement learning process for mimo-v2.6-pro and mimo-v2.6-flash. The pro run (step 15) has cost $1,059,782 and processed 30.2B tokens; the flash run (step 20) cost $481,213 and processed 46.2B tokens. Benchmarks: DeepSWE v1.1 mini-swe-agent avg@3 scores are 65.78 (pro) and 60.77 (flash). Notices report several restarts due to VRAM issues, network connectivity, and infra errors; the cyber dataset was removed from the upcoming pro run due to bad patterns in rollout logs.

Code datasets make up ~68% of training prompts for the pro run. The dashboard includes per-source acceptance logs, step timing, and dynamic sampler statistics.

What commenters are saying

Commenters largely praise the transparency as remarkable in an otherwise secretive industry, though some question why Xiaomi would publish such detail. Several users report positive hands-on experience with Mimo 2.5, citing strong ROI and good instruction-following for coding. A thread debates whether running benchmarks during training constitutes contamination; consensus holds that using benchmarks as stopping criteria is common and distinct from training on them. One commenter notes China's official policy favoring open models may incentivize this openness.