Skip to content
The Eval Room

Article

Beam Announced. What Are the Controls Actually Holding Back?

Beam is announced but not out. Chinese rivals lead its own table. A White House accusation and a GPU smuggling indictment raise a question neither answers.

Reflection AI announced Beam on October 5, 2026: a 501-billion-parameter open-weight model for coding, reasoning, and agentic work, with weights promised under Apache 2.0 later this month. Reflection says it is giving selected users early access while public weights remain scheduled for later this month. Beam's own table puts frontier Chinese open models ahead on raw capability; a White House official has accused the maker of one of them of accessing GB300s abroad; and, separately, a man has been indicted over alleged smuggling of GPU servers to China. Neither the accusation nor the indictment shows what any model in Beam's table was trained on.

What Beam is

Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, of which only 23 billion are active per token. That architecture gives the model large total capacity while keeping per-token compute lower than a dense model of comparable quality. Reflection says it pretrained on 23.8 trillion tokens, then ran a four-week reinforcement learning campaign on 10.5K Nvidia GB300 GPUs generating more than 100 million training rollouts.

The model is text-only and does not support multimodal inputs, according to TechCrunch. Reflection's planned release covers the weights, a technical report, a model card, and tools for running, evaluating, and fine-tuning the model. Training code and pretraining data are not listed among what will be released.

Beam's training at a glance

Total / active parameters
501B / 23B
Pretraining tokens
23.8T
RL rollouts generated
100M+
GB300 GPUs in the RL run
10.5K
Source: reflection.ai

What the benchmarks claim

Reflection's announcement carries a full comparison table alongside individual scores: AIME 2026 at 97.8%, GPQA Diamond at 90.5%, Terminal Bench v2.1 at 80.1%, and SWEBench Verified at 80.9%. No independent results have been published — the scores appear only in the vendor's own post. TechCrunch's coverage references general performance claims but does not reproduce the specific numbers.

Reflection's own framing is direct: its announcement states Beam "advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks," while acknowledging that "frontier open models like Kimi K3 remain ahead on raw capability." The company's own benchmark table shows this gap on Terminal Bench v2.1, where Beam scores 80.1 against GLM 5.2 (81.0), GLM 5.3 (88.2), Kimi K3 (88.3), Qwen 3.8-Max (86.6), and DeepSeek V4.1 Flash (90.6); on DeepSWE v1.1, Beam's 44.4 trails Kimi K3 (68.0) and DeepSeek V4.1 Flash (74.2). The gap extends to the reasoning benchmarks: on AIME 2026, GLM 5.2 scores 99.2 against Beam's 97.8; on GPQA Diamond, GLM 5.2 (91.2), GLM 5.3 (91.7), Kimi K3 (93.5), Qwen 3.8-Max (92.6), and DeepSeek V4.1 Flash (90.9) all lead Beam's 90.5. Beam leads GLM 5.2 on SWE Bench Pro v1 (65.5 vs 62.1), though Qwen 3.8-Max (67.7) leads Beam on that benchmark; on the coding, reasoning, and tool-calling tabs it leads both Inkling and Nemotron 3 Ultra on every row with a reported score; on the general-capabilities tab it trails both on IFBench (Beam 79.7, Inkling 79.8, Nemotron 3 Ultra 81.7) and trails Inkling on AA Omniscience (Beam 13.0, Inkling 14.2). All figures are from the vendor's own announcement.

The pitch is not raw capability but inference efficiency. Reflection claims 3–4x less inference compute than comparable models, estimating generation forward-pass FLOPs as 2 × active parameter count × mean generated tokens per attempt, counting each multiply-add as two operations — and using activated parameters, not total model size, for MoE models. The figure excludes prompt prefill, context-dependent attention operations, and serving overhead; Reflection's own description calls it "an approximate compute comparison rather than measured inference cost."

Beam's score versus estimated generation FLOPs across three tasks (DeepSWE v1.1, HLE text-only, Terminal Bench 2.1), compared with Inkling, Nemotron 3 Ultra, Muse Glimmer, GLM-5.2, and Qwen3.8.
Image: Reflection AI

Who built Reflection AI

Reflection AI was reportedly founded in March 2024 by Misha Laskin and Ioannis Antonoglou, both former Google DeepMind researchers; Antonoglou was a co-creator of AlphaGo. Beam is Reflection's first open-weight model. TechCrunch, citing PitchBook, reports the company has raised roughly $4.7 billion, with backers including Nvidia, Sequoia Capital, and Lightspeed Venture Partners.

TechCrunch reports Reflection is aiming Beam and future models at enterprises and sovereign nations, with a pitch to build 'AI factories' — a product that would let institutions build their own customized, local AI system by training Reflection's AI models on their own proprietary data.

Open-weight is not open-source

Reflection describes Beam as "open-weight," not open-source, and the distinction matters. With the weights alone, you can download, run, and fine-tune the model — but you do not have the training code or pretraining data that would let you rebuild it. The announcement does not list training code or pretraining data among what will be released.

As of October 5, 2026, the public weights are not yet available, though Reflection says it is making an early version available to a select group of users. The announcement reads: "This month, we will release the weights under an Apache 2.0 license, along with documentation and the full stack for running, evaluating, and fine-tuning the model." Until independent results are published, every benchmark claim rests on Reflection's own reporting alone.

Beam's safety evaluation is also not public at announcement time. Reflection states it has applied deliberative alignment techniques and adversarial dataset construction, but has not published red-team results or an independent audit report.

What the Export Cases Establish, and What They Don't

Reflection's announcement states Beam was pretrained on 6,144 NVIDIA GB300 NVL72 GPUs and that its RL campaign ran on 10.5K GB300 GPUs. The January 2026 BIS rule moves license review from a presumption of denial to case-by-case only for exports from the United States to end users in China or Macau, of chips with total processing performance below 21,000 and DRAM bandwidth below 6,500 GB/s — BIS cites Nvidia's H200 and AMD's MI325X as examples — and only when the applicant meets the rule's certification conditions. Reexports, including exports from abroad, and in-country transfers of those same chips to or within China or Macau keep a presumption of denial, as a Morgan Lewis analysis notes; chips above those thresholds remain subject to a presumption of denial. Neither the rule nor the Morgan Lewis analysis mentions Blackwell or the GB300. Reuters reported in February 2026 that export controls 'currently bar Blackwell shipments to China,' and quoted a senior Trump administration official as saying 'we're not shipping Blackwells to China.'

Bloomberg Law reported July 22, 2026 that White House OSTP Director Michael Kratsios made a social media post stating Moonshot AI 'acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models,' a post Bloomberg Law described as 'making the case that Moonshot violated US export control rules and companies' terms of service to build its product.' In the same article, Bloomberg Law reported that 'a White House official accused China's Moonshot of improperly using US artificial intelligence models and Nvidia Corp. chips to create the Kimi K3 system.' Reuters, in its July 31, 2026 report, noted that Alibaba denied supplying H200 chips to Moonshot but did not identify the actual chip type, and that Moonshot 'could not be reached for comment outside regular business hours.' No public response from Moonshot to the chip allegations appeared in the public record as of October 5, 2026.

A separate case: on October 1, 2026, the Justice Department announced the indictment of Greg Lui, 38, owner of Earthmade Computer Inc., on three counts — conspiracy to violate the Export Control Reform Act and the Export Administration Regulations, outbound smuggling, and conspiracy to commit money laundering. The allegation: Lui and co-conspirators shipped 'more than $300 million' in export-controlled computer servers to China via Malaysia and Singapore. FBI Assistant Director Roman Rozhavsky stated the investigation 'revealed that Lui allegedly sold the Chinese government hundreds of millions of dollars' worth of American Super Intelligence technology.' No Chinese AI company or model is named in the Justice Department's announcement of the charges. The Justice Department's own notice: 'An indictment is merely an allegation. All defendants are presumed innocent until proven guilty beyond a reasonable doubt in a court of law.'

Training hardware for all five Chinese models in Beam's benchmark table — Kimi K3, GLM 5.2, GLM 5.3, Qwen 3.8-Max, and DeepSeek V4.1 Flash — is not publicly confirmed. Neither allegation proves how any Chinese model in Beam's table was trained: the accusation carries the word 'likely' and has not been answered with confirmation, and the Justice Department's announcement names no AI company or model.

What are the controls actually holding back?

Sources