Reflection's Beam Is the US Open-Weight Coding Model Aimed at DeepSeek and Qwen

Reflection's Beam Is the US Open-Weight Coding Model Aimed at DeepSeek and Qwen

6 min readOctober 7, 2026

Quick verdict

Reflection announced Beam, a text-only mixture-of-experts model with 501B total parameters and 23B active, built for coding, agentic, and scientific work. The company says it trained from scratch and will release full weights under Apache 2.0 this month. The headline claims are 80.9 on SWE-bench Verified and three to four times the inference efficiency of GLM-5.2. The reason this matters is not the single benchmark. It is that a US lab is shipping an open-weight model squarely aimed at DeepSeek and Qwen, the two names that have owned the open frontier all year. Early independent reads put Beam around GLM-5.2 level and still behind the top Chinese open models, so treat the launch as a serious entrant rather than a new leader until the weights and the tech report are out.

What actually shipped

The announcement came from Reflection and co-founder Misha Laskin. Beam is a 501B-total, 23B-active MoE, text only for now, with full weights promised under Apache 2.0 within the month and a tech report plus open-source integrations to follow, per Alex Polozov. So the thing people can actually download does not exist yet. What landed is the announcement, a set of claimed numbers, and early access for a few independent evaluators.

The training story is where the scale shows. The team's data lead says Beam saw 23.8 trillion pretraining tokens, a chunk of it pulled through an OCR pipeline over hundreds of millions of PDFs (source). Brandon Damos describes a stable reinforcement-learning run on roughly 10,000 GB300 GPUs, with more than 100 million rollouts across about a million tasks (post). A widely shared summary of Reflection's claims lists 80.9 on SWE-bench Verified, the 3-4x inference-efficiency edge over GLM-5.2, and four weeks each of pretraining and RL on about 10,500 GB300s.

The benchmark and the caveats

An 80.9 on SWE-bench Verified would put Beam in real coding-model territory, close to where the strong closed models were not long ago. But the number is first-party, and the people with early access are more measured. Artificial Analysis has early access and expects Beam to be among the most token-efficient open models for its intelligence level, which is a real compliment but not the same as a top score. One critique places it below DeepSeek V4 Flash on some benchmarks, and Nathan Lambert groups Beam with Nvidia's and Thinking Machines' releases as strong US open models that still trail their Chinese counterparts. The honest read: good, efficient, not clearly ahead.

The architecture analysis points the same way. Elie Bakouch estimates only about 12% BF16 model-FLOPs-utilization in pretraining and reads the design as a 3:1 interleaved global and sliding-window attention, with better held-out code perplexity than DeepSeek V4. Teortaxes goes further and calls Beam essentially an iso-FLOP replication of DeepSeek V3, inferring something like 1.3 billion RL sandboxes over the four-week run. In plain terms, a lot of this is proven technique executed at scale by a well-funded US team, not a new architecture that resets the field.

Why a US open model is the real story

Open weights in 2026 have mostly meant Chinese labs: DeepSeek, Qwen, GLM, Kimi. A reddit thread flagged the Axios report on Beam under the frame "a US open-weight model to take on DeepSeek and Qwen," and that framing is why the launch drew attention past its benchmark. Andrew Curran's pre-launch reporting cited Axios figures of roughly $150M a month for Colossus compute plus a $1B Nebius deal, and said other unnamed US labs plan open releases this month too. If that holds, the open-weight conversation stops being a one-country story, and enterprises that want to self-host get US-made options with Apache 2.0 terms.

The branding carries baggage, though. Commenters were quick to recall the Reflection 70B episode from two years ago, where benchmark claims did not survive scrutiny. That history is exactly why the "we will release the weights and a tech report" promise matters more than the slide of numbers. Until the weights are public and third parties run their own evals, Beam is a credible announcement, not a settled result.

Video: Reflection Beam explained

A walk through what Reflection announced, the training scale behind Beam, and how it stacks up against the Chinese open models it is chasing.

What it means if you pay for AI

For most people the practical answer is: nothing yet, and then possibly a lot. You cannot run a 501B-total model on a laptop, so "open weights" here means cheap hosted access through providers once the weights drop, not local inference. If the coding scores hold up on independent evals, Beam becomes another model you route the bulk of your coding and agent work to while keeping a closed frontier model for the genuinely hard reasoning. That model-by-task split is the whole case for running everything through one aggregator instead of paying for a stack of separate subscriptions. For now, watch for the weights and a third-party SWE-bench run before you move any real workload.

FAQ

Can I download and run Beam today?

Not yet. Reflection announced the model and says full Apache 2.0 weights are due this month. Until they land, only a handful of independent evaluators have early access, and even then a 501B-total model realistically means renting it from a provider rather than self-hosting. See our guide to the best AI coding agents for where a model like this fits.

Is 80.9 on SWE-bench Verified real?

It is Reflection's own claimed number, and independent reviewers with early access are more cautious, placing Beam around GLM-5.2 level and below DeepSeek V4 Flash on some tests. Treat the 80.9 as a first-party figure until third parties publish their own runs.

How is Beam different from DeepSeek and Qwen?

The headline difference is origin and license: Beam is a US-built model shipping under Apache 2.0, aimed at the open-weight niche the Chinese labs have dominated. On architecture it looks like a scaled, efficient take on proven MoE designs rather than something fundamentally new.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles