Models

Reflection's Beam: A 501B Open-Weight Model Built for Coding

Reflection AI has introduced Beam, a 501-billion-parameter open-weight model built for coding and agentic tasks. The weights come later this month under Apache 2.0, but the benchmark scores are still Reflection's own.

Daniel HarrisDaniel Harris
Heat: 1,250
Reflection's Beam: A 501B Open-Weight Model Built for Coding

Reflection AI's first open-weight model packs 501 billion parameters and lands close to far bigger Chinese systems on coding tests. The weights, technical report, and fine-tuning stack are promised for later this month.

What Reflection Beam Actually Is

Reflection AI, a US lab, introduced Beam on October 5. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, of which 23 billion are active per token. That design keeps the model large on paper while only waking part of it for each word it produces, so running it costs less than the parameter count suggests.

The company trained Beam with a heavy focus on coding, reasoning, and agentic workloads, meaning jobs where the model plans steps, calls tools, and works through a task on its own rather than answering a single question. It's text-only, so it won't handle images or audio.

Beam isn't downloadable yet. Reflection is running final red-teaming and evaluations and has opened early access sign-ups. Weights release under Apache 2.0, along with the technical report, model card, and developer artifacts, is slated for later this month. That license matters to businesses. Apache 2.0 lets them run and modify the model commercially without the usage strings attached to some open releases.

How Beam Was Trained: 23.8 Trillion Tokens and 100M Rollouts

Reflection pretrained Beam on 23.8 trillion tokens drawn from the web and licensed datasets, a figure in line with similar-sized open base models. That's the raw reading diet that gives the model its general knowledge.

The more unusual part is the reinforcement learning stage. Reflection ran the RL phase on 10,500 NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256,000 tokens. The company says this is one of the largest RL runs by any open lab to date. RL is the step where a model practices problems, gets graded, and adjusts, so a bigger run usually means better multi-step reasoning.

For readers wondering why that matters: heavy RL tends to show up in tasks where the model has to try, check, and correct itself, like debugging code or tracing a bug across several files. Reflection notes that its scores kept climbing as RL compute grew.

Beam's Coding Benchmarks, According to Reflection

All numbers here are self-reported by Reflection, not independently verified. On SWE-bench Verified, which measures whether a model can fix real GitHub issues, Beam scored 80.9. On Terminal Bench v2.1, a test of command-line and multi-step terminal tasks, it hit 80.1.

Those are strong numbers, but they sit below several Chinese open models. Kimi K3 reported 88.3 on Terminal Bench, DeepSeek V4.1 Flash 90.6, Qwen 3.8-Max 86.6, and GLM 5.3 88.2. Against GLM 5.2, Beam is close, at 81.0 versus its own 80.1.

Reflection is upfront about the gap. The company says Beam is competitive with GLM 5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, while models like Kimi K3 remain ahead on raw capability.

Model

Terminal Bench v2.1

SWE-bench Verified

Reflection Beam

80.1

80.9

GLM 5.2

81.0

not reported

GLM 5.3

88.2

not reported

Kimi K3

88.3

not reported

Qwen 3.8-Max

86.6

not reported

Beam's pitch isn't winning the leaderboard. It's getting close to the leaders while costing less to run.

The Real Selling Point: Inference Efficiency

Where does Beam actually win? On efficiency at inference time, the moment a model is answering a request. Reflection says Beam reaches reasoning scores comparable to GLM 5.2 while using 3 to 4 times less inference compute, and that the gap widens against models in the 2-trillion-parameter family like Qwen 3.8-Max.

Treat that claim carefully. The comparison isn't a measured dollar cost. Reflection estimated forward-pass compute using an approximation built on active parameter counts and generated tokens, drawing on data from Artificial Analysis and DataCurve. The estimate excludes prompt processing, attention, and serving overhead, which the company acknowledges makes it an approximate comparison rather than a real cost benchmark.

Still, the underlying logic is sound. A sparse model that activates 23 billion parameters per token should draw less compute than a dense model of similar ability. For teams paying per token, that can translate into lower bills on high-volume coding and agent workflows.

What to Watch Before the Weights Drop

The biggest open question is verification. Every score so far comes from Reflection's own suite, and the broader public hasn't run Beam yet. Independent testing after release will settle whether the efficiency claim holds in production.

A promise, not proof.

Availability is the other factor. With weights promised within weeks under a permissive license, teams weighing open models against API-only options have a reason to hold off on long-term commitments until they can test Beam themselves. If the efficiency numbers survive scrutiny, a 501B model that runs like something smaller is a genuinely useful option for enterprise coding work.

Share This Story

Sources

Related AI News

GPT-6 and Intelligent UI Roll Out to All ChatGPT Users
Models

GPT-6 and Intelligent UI Roll Out to All ChatGPT Users

ChatGPT's answers are turning into buttons, charts, and mini-apps. GPT-6 and the Intelligent UI behind them are now reaching everyone, free users included.

Heat: 1,700
Anthropic's Haiku 5.5 Cuts Small-Model API Costs 75%
Models

Anthropic's Haiku 5.5 Cuts Small-Model API Costs 75%

Anthropic's cheapest model yet lands at roughly a quarter of its predecessor's running cost. Here's what Claude Haiku 5.5 means for teams that bill by the token.

Heat: 1,500
Mistral Large 4 Preview: Europe's 1T Open-Weight Bet
Models

Mistral Large 4 Preview: Europe's 1T Open-Weight Bet

Mistral says its new flagship is the strongest open-weight model to come out of Europe or the US. It's a public preview for now, and the weights don't land until the end of the month, so what you can actually test today is the API.

Heat: 1,600
Nano Banana 2.1 Is Out: Half the Price, 23 Days to Switch
Models

Nano Banana 2.1 Is Out: Half the Price, 23 Days to Switch

Google's newest image model costs about half of what the last one did, and it drops more than just the per-image price. Anyone still calling Nano Banana 2 through the API has 23 days to move.

Heat: 1,300