Reflection Beam is the first frontier-scale open-weight model from a US startup that can credibly sit at the same table as China’s best, and Reflection AI’s own launch post is surprisingly honest about where it loses. Beam is a 501 billion parameter mixture-of-experts model with 23 billion parameters active per token, announced on October 5, 2026, with Apache 2.0 weights promised later this month.
I read Reflection’s launch post line by line, including the four benchmark tables that most of the coverage summarised as “rivals Chinese models”. The tables tell a more specific story. Beam is competitive with the open models of mid-2026. It is not competitive with the ones that shipped in the last few weeks.
What Is Reflection Beam?
Reflection Beam is a text-only sparse mixture-of-experts model with 501B total and 23B active parameters, trained by Reflection AI for coding, reasoning and agent work. Reflection was founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. According to TechCrunch, it has raised about $4.7 billion and is valued at $25 billion, with Nvidia among its backers.
The specs from Reflection’s launch post:
- Parameters: 501B total, 23B active per token
- Pretraining: 23.8 trillion tokens, on 6,144 Nvidia GB300 GPUs, in under four weeks
- Reinforcement learning: 10,500 GB300 GPUs for four weeks, more than 100 million rollouts, roughly 1.3 billion sandboxes
- Context: pretrained at up to 256K, extended to an effective 1M tokens in midtraining
- License: Apache 2.0, with weights, technical report and model card “later this month”
- Access today: early-access waitlist at platform.reflection.ai
The RL numbers are the part I keep coming back to. A million coding, agentic and STEM environments, with up to 170,000 sandboxes running at once, is the kind of infrastructure only a handful of labs have built. That is where the money went, and it shows in the reasoning scores.
Beam is a serious model from a serious lab, but you cannot download it yet, and that matters more than any number below.
Reflection Beam Benchmarks: The Coding Table
Reflection published scores against seven open models: Thinking Machines’ Inkling, Nvidia’s Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. Here are the agentic coding rows for the three models that matter most to a developer picking an open model in October 2026.
| Benchmark | Reflection Beam | Kimi K3 | GLM 5.3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|
| Terminal-Bench v2.1 | 80.1 | 88.3 | 88.2 | 90.6 |
| SWE-Bench Pro v2-Hard | 77.2 | 88.2 | 84.3 | — |
| DeepSWE v1.1 | 44.4 | 68.0 | 61.0 | 74.2 |
| SWE Atlas Codebase QnA | 34.6 | 68.0 | 61.0 | — |
(Source: Reflection AI, “Introducing Beam”, October 5, 2026. All scores vendor-reported; ”—” means Reflection did not report a score.)
Look at the last two rows. A 44.4 on DeepSWE against 74.2 for DeepSeek V4.1 Flash, a model with “Flash” in its name, is not a small gap. On SWE Atlas Codebase QnA, which tests whether a model can answer questions about a large repository across a long interaction, Beam scores roughly half of what Kimi K3 does.
That second number is the one I would worry about. Answering questions about a codebase you have been working in for an hour is most of what a coding agent actually does all day.
Here is the fair reading. Reflection compares Beam most directly with GLM 5.2, and against that model it holds up: 65.5 vs 62.1 on SWE-Bench Pro v1, and 80.1 vs 81.0 on Terminal-Bench v2.1. It also clearly beats Inkling and Nemotron 3 Ultra, the two Western open models in the table, on most coding rows.
On Reflection’s own data, Beam beats the Western open models and roughly ties GLM 5.2, but it trails the current Chinese leaders by 7 to 33 points on agentic coding.
Where Beam Actually Wins: Efficiency and Reasoning
The headline Reflection chose is compute, not raw score. According to TechCrunch’s launch coverage, Reflection says Beam matches GLM 5.2 on advanced reasoning while using 3 to 4 times less inference compute. Beam activates 23B parameters per token, and Reflection’s chart counts both reasoning and answer tokens.
Reflection’s own footnote limits that claim. The estimates “exclude prompt prefill, context-dependent attention operations, and serving overhead”. In other words, it is an arithmetic comparison of active parameters times generated tokens, not a measured bill. For coding agents, where prefill on a 100K-token context is a large share of the cost, that exclusion is not a detail.
The reasoning numbers are better than the coding ones:
| Benchmark | Reflection Beam | Kimi K3 | GLM 5.3 | Inkling |
|---|---|---|---|---|
| AIME 2026 | 97.8 | — | — | 97.1 |
| GPQA Diamond | 90.5 | 93.5 | 91.7 | 87.2 |
| HLE (no tools) | 36.2 | 46.9 | 42.3 | 29.7 |
| MCP Atlas | 78.7 | 82.3 | 84.2 | 76.0 |
(Source: Reflection AI launch post, October 5, 2026. Vendor-reported.)
Beam is within 3 points of the leaders on GPQA Diamond and posts the top AIME score in the set. HLE is a 10-point gap. MCP Atlas, which measures tool calling across MCP servers, is a 5.5-point gap behind GLM 5.3. If you have been following the MCP servers guide, that is the benchmark closest to how agents use tools in practice.
Beam’s real strength is reasoning per unit of compute, which matters to whoever pays for the GPUs, more than to whoever picks the model in their editor.
Reflection Beam vs Kimi K3 vs GLM 5.3: Which Open Model for Coding?
For a developer choosing an open model for agentic coding today, the decision is short.
If you want the strongest open coding model you can call this week, Reflection’s own table points to Kimi K3 or GLM 5.3. Both are available now on hosted inference, and as I cover in the companion post on running Claude Code with open models through Together Link, Together AI lists Kimi K3 at $2.70 / $13.50 per million input/output tokens and GLM 5.3 at $1.40 / $4.40.
If you want the cheapest model that still clears a high bar, DeepSeek V4.1 Flash posts the best Terminal-Bench and DeepSWE numbers in Reflection’s table, at $0.30 / $1.20 on Together AI.
Beam makes sense for a narrower group:
- Teams that need a US-trained model for procurement, compliance or government work, where a Chinese-origin model is a non-starter regardless of score
- Teams that self-host and care about active parameters per token, because 23B active is light for a model of this size
- Fine-tuners who want an Apache 2.0 base with strong reasoning to specialise on their own codebase
That first group is larger than benchmark watchers tend to assume. The TechCrunch and The Hill coverage both frame Beam as an answer to Chinese open models, and for a lot of US enterprises, that framing is the whole purchase decision.
For most individual developers, Kimi K3 and GLM 5.3 are the better coding models today; Beam is the better choice when the model’s origin and license matter as much as its score.
What Running Beam Will Actually Take
Once the weights land, the 23B active figure will tempt people into thinking Beam runs on a workstation. It will not.
Mixture-of-experts models still need all 501 billion parameters in memory, because any expert can be selected for any token. At 8-bit precision that is roughly 500GB of weights before you add KV cache for long contexts. At 4-bit it is still around 250GB. This is a multi-GPU server model, not a laptop model.
What the sparse design does buy you is throughput. Each token only touches 23B parameters, so a node that can hold the model serves it much faster than a dense model of similar total size.
I would also expect quantised community builds soon after the weights ship. Plan on testing those against your own tasks rather than trusting the launch table, because quantisation tends to hurt long-horizon agentic work first, which is exactly where Beam is already weakest.
Budget for a multi-GPU node to self-host Beam, and expect hosted providers to be the practical option for most teams.
The Caveats Before You Plan Around Beam
Four things are missing as of October 6, 2026, and each one should stop you from making Beam a dependency this month.
There are no weights. Reflection says “later this month”. Open-weight launches slip, and a model still in “final red-teaming and evaluations” can change between preview and release.
There is no price. The launch post does not list API pricing, and the early-access API is waitlisted.
There is no technical report. It is promised alongside the weights. Until it ships, details like the expert count, attention design and data mix are unknown.
There are no independent benchmarks. Every number in this post comes from Reflection. TechCrunch notes the claims have not been independently verified, and I would wait for Artificial Analysis or a similar independent run before trusting any lead under 5 points. The same rule applied to Gemini 4 Argon last week, which is also announced but not generally available.
What I respect about this launch is that Reflection published the rows where Beam loses badly. Plenty of labs would have dropped DeepSWE and SWE Atlas from the table. A launch that shows a 34.6 next to a 68.0 is easier to trust on everything else.
Treat Beam as a promising preview, then re-evaluate when the weights, technical report and independent scores arrive later in October.
Should Developers Care About Reflection Beam?
Yes, but not this week, and not for the reason the headlines give.
The story is not that a US startup beat China. On Reflection’s own numbers, it did not, at least against Kimi K3, GLM 5.3 and DeepSeek V4.1 Flash. The story is that a US lab now has the infrastructure to train a 501B model with a million RL environments, and it is releasing the result under Apache 2.0. The other Western open models in Reflection’s table are Inkling and Nemotron 3 Ultra, and Beam beats both on every coding row where both scores are reported.
That gives American enterprises, and anyone else with a reason to avoid Chinese-origin weights, a real open option for the first time. It also gives fine-tuners a strong base to build on.
For day-to-day coding with an agent like Claude Code or Codex, nothing changes yet. The Cursor vs Claude Code vs Codex comparison still applies to most developers, because the choice of agent harness matters more than which open model you plug into it.
My verdict: watch for the weights and independent benchmarks, and keep using Kimi K3, GLM 5.3 or DeepSeek V4.1 Flash if you need an open coding model today.