Beam: Reflection's 501B open-weight model
Points and comments are a snapshot, not live.
Reflection announces Beam, a 501B open-weight MoE model for coding and agentic tasks.
Beam is a sparse Mixture-of-Experts model with 501B total parameters (23B active), pretrained on 23.8 trillion tokens. A high-compute RL run used 10.5K NVIDIA GB300 GPUs over four weeks, generating over 100 million rollouts. The model focuses on coding, reasoning, and agentic capabilities, claiming competitive performance with frontier open-weight models on benchmarks like DeepSWE and SWE-Bench, while offering 3-4× inference compute efficiency. Reflection will release weights, a technical report, and developer artifacts later this month.
The pretraining recipe was designed for stable MoE optimization, using auxiliary-loss-free load balancing with cosine decay. The RL infrastructure supported up to 170K concurrent sandboxes across two clouds and four regions, with 71 inference incidents handled without terminating training. A controllable length penalty allows users to adjust reasoning effort against token usage.
What commenters are saying
Skepticism dominates: commenters question the value of announcing a model without releasing weights or providing access beyond a sign-up form. Some see it as a promotional tactic rather than a genuine open-weight contribution, recalling past broken promises of releases. A few defend the need for promotion when companies give away valuable model weights, but the thread is largely critical.
Specific criticism targets the claim that a "world map" puzzle is too recent to appear in training data; commenters note the puzzle existed since August 2025, making the timing argument sloppy. Others debate data sourcing: some defend proprietary data as necessary for quality, while others dismiss it as likely unremarkable. The need for great data over modeling tricks is a recurring theme.