Reflection AI's first open-weight model packs 501 billion parameters and lands close to far bigger Chinese systems on coding tests. The weights, technical report, and fine-tuning stack are promised for later this month.
What Reflection Beam Actually Is
Reflection AI, a US lab, introduced Beam on October 5. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, of which 23 billion are active per token. That design keeps the model large on paper while only waking part of it for each word it produces, so running it costs less than the parameter count suggests.
The company trained Beam with a heavy focus on coding, reasoning, and agentic workloads, meaning jobs where the model plans steps, calls tools, and works through a task on its own rather than answering a single question. It's text-only, so it won't handle images or audio.
Beam isn't downloadable yet. Reflection is running final red-teaming and evaluations and has opened early access sign-ups. Weights release under Apache 2.0, along with the technical report, model card, and developer artifacts, is slated for later this month. That license matters to businesses. Apache 2.0 lets them run and modify the model commercially without the usage strings attached to some open releases.
How Beam Was Trained: 23.8 Trillion Tokens and 100M Rollouts
Reflection pretrained Beam on 23.8 trillion tokens drawn from the web and licensed datasets, a figure in line with similar-sized open base models. That's the raw reading diet that gives the model its general knowledge.
The more unusual part is the reinforcement learning stage. Reflection ran the RL phase on 10,500 NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256,000 tokens. The company says this is one of the largest RL runs by any open lab to date. RL is the step where a model practices problems, gets graded, and adjusts, so a bigger run usually means better multi-step reasoning.
For readers wondering why that matters: heavy RL tends to show up in tasks where the model has to try, check, and correct itself, like debugging code or tracing a bug across several files. Reflection notes that its scores kept climbing as RL compute grew.
Beam's Coding Benchmarks, According to Reflection
All numbers here are self-reported by Reflection, not independently verified. On SWE-bench Verified, which measures whether a model can fix real GitHub issues, Beam scored 80.9. On Terminal Bench v2.1, a test of command-line and multi-step terminal tasks, it hit 80.1.
Those are strong numbers, but they sit below several Chinese open models. Kimi K3 reported 88.3 on Terminal Bench, DeepSeek V4.1 Flash 90.6, Qwen 3.8-Max 86.6, and GLM 5.3 88.2. Against GLM 5.2, Beam is close, at 81.0 versus its own 80.1.
Reflection is upfront about the gap. The company says Beam is competitive with GLM 5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, while models like Kimi K3 remain ahead on raw capability.
Model | Terminal Bench v2.1 | SWE-bench Verified |
|---|---|---|
Reflection Beam | 80.1 | 80.9 |
GLM 5.2 | 81.0 | not reported |
GLM 5.3 | 88.2 | not reported |
Kimi K3 | 88.3 | not reported |
Qwen 3.8-Max | 86.6 | not reported |
Beam's pitch isn't winning the leaderboard. It's getting close to the leaders while costing less to run.
The Real Selling Point: Inference Efficiency
Where does Beam actually win? On efficiency at inference time, the moment a model is answering a request. Reflection says Beam reaches reasoning scores comparable to GLM 5.2 while using 3 to 4 times less inference compute, and that the gap widens against models in the 2-trillion-parameter family like Qwen 3.8-Max.
Treat that claim carefully. The comparison isn't a measured dollar cost. Reflection estimated forward-pass compute using an approximation built on active parameter counts and generated tokens, drawing on data from Artificial Analysis and DataCurve. The estimate excludes prompt processing, attention, and serving overhead, which the company acknowledges makes it an approximate comparison rather than a real cost benchmark.
Still, the underlying logic is sound. A sparse model that activates 23 billion parameters per token should draw less compute than a dense model of similar ability. For teams paying per token, that can translate into lower bills on high-volume coding and agent workflows.
What to Watch Before the Weights Drop
The biggest open question is verification. Every score so far comes from Reflection's own suite, and the broader public hasn't run Beam yet. Independent testing after release will settle whether the efficiency claim holds in production.
A promise, not proof.
Availability is the other factor. With weights promised within weeks under a permissive license, teams weighing open models against API-only options have a reason to hold off on long-term commitments until they can test Beam themselves. If the efficiency numbers survive scrutiny, a 501B model that runs like something smaller is a genuinely useful option for enterprise coding work.






