█

AI/TLDR

Beam

Reflection AI's first open-weight model, announced October 5, 2026: a 501B-total, 23B-active sparse Mixture-of-Experts for coding, reasoning and agents, with a 1M-token context and Apache 2.0 weights promised for October 2026.

BeamAPI onlyPreview — early access for a select group of users through the waitlist at platform.reflection.ai; Reflection AI says the weights, technical report, model card and developer artifacts will be released later in October 2026.
Released
5 Oct 2026
Context
1M tokens (effective context, extended in midtraining)
Parameters
501B total · 23B active
License
Apache-2.0 (announced — Reflection AI says it will release the weights under Apache 2.0 in October 2026)
Coverage
1 story

Overview

Beam is Reflection AI's first open-weight model, announced on October 5, 2026. It is a sparse Mixture-of-Experts language model with 501 billion total parameters, of which 23 billion are active per token, built for coding, reasoning and agentic work. Reflection AI describes Beam as the first model in a series.

Reflection AI chart of score against estimated generation forward FLOPs per attempt on DeepSWE v1.1, HLE (text-only) and Terminal Bench 2.1, with Beam's reasoning-effort curve reaching higher scores at lower compute than Inkling, Nemotron 3 Ultra, Muse Glimmer and GLM-5.2, and Qwen3.8 points at far higher compute.
Beam's score against generation compute across reasoning efforts, from Reflection AI's launch post.Reflection AI ↗

Beam is a text model; Reflection AI says it can work with other modalities when they are represented as text. Midtraining extends its effective context length to 1M tokens, and its reinforcement-learning stage used a maximum context of 256K tokens.

Reflection AI pretrained Beam on 23.8 trillion tokens in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. Its reinforcement-learning run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks of training.

Reflection AI line charts of Beam's DeepSWE, HLE and Terminal Bench 2.1 scores rising as the number of reinforcement-learning rollouts grows from near zero to about 80 million.
Scaling the rollout budget: Beam's scores against the number of RL rollouts.Reflection AI ↗

In Reflection AI's launch tables Beam scores 80.9 on SWE-bench Verified, 80.1 on Terminal Bench v2.1, 78.0 on SWE-bench Multilingual, 97.8 on AIME 2026 and 90.5 on GPQA Diamond, set against Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. Reflection AI updated Beam's results on October 8, 2026; the table on this page uses the updated launch-post figures.

At announcement Beam was in final red-teaming and evaluation, open to a select group of users through a waitlist on Reflection AI's platform. Reflection AI says it will release the weights under the Apache 2.0 license later in October 2026, together with a technical report, a model card and developer artifacts.

Released2026-10-05
LicenseApache-2.0 (announced — Reflection AI says it will release the weights under Apache 2.0 in October 2026)
WeightsAPI only
Parameters501B total · 23B active
Context1M tokens (effective context, extended in midtraining)
ArchitectureSparse Mixture-of-Experts
ModalitiesText
StatusPreview — early access for a select group of users through the waitlist at platform.reflection.ai; Reflection AI says the weights, technical report, model card and developer artifacts will be released later in October 2026.

Benchmarks

Beam against open models, transcribed from Reflection AI's launch post (results updated October 8, 2026). Blank cells were not reported; Reflection AI used Artificial Analysis and DataCurve as sources for other models' evals.

BenchmarkBeamInklingNemotron 3 UltraGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 Flash
DeepSWE v1.144.4——4461685174.2
SWE Bench Pro v2-Hard77.256.9——84.388.2——
SWE Bench Pro v165.554.346.462.1——67.7—
Terminal Bench v2.180.163.856.48188.288.386.690.6
SWE Atlas Codebase QnA34.6———6168——
SWE-bench Multilingual78—67.7—————
SWE-bench Verified80.977.670.7—————
AIME 202697.897.1—99.2————
HLE (no tools)36.229.726.740.542.346.943.639.1
SciCode49.746.144.6—5958.752.152
CritPT (AA)16.35.43.120.919.123.42014.3
GPQA Diamond90.587.28791.291.793.592.690.9
AutomationBench (public)37——26.248.246.739.854.8
MCP Atlas78.77663.177.884.282.384.5—
tau3 banking382522.637.1—37.155.2—
BrowseComp (w/ context management)77.477.144.4——91.2——
DeepSearchQA (w/ context management)80.1————95——
AA-LCR79.377.379.378.379.788.780.384
LongBench v265.5—61.964——66.3—
IFBench79.779.881.773.3——82.8—
AA Omniscience Index (public split, Reflection's runs)1314.28.6—22.6——6.6

Comparison source ↗

This model's scores

  1. AIME 202697.8%
  2. GPQA Diamond90.5%
  3. SWE-bench Verified80.9%
  4. Terminal Bench v2.180.1%
  5. IFBench79.7%
  6. AA-LCR79.3%
  7. MCP Atlas78.7%
  8. SWE-bench Multilingual78%
  9. BrowseComp (with context management)77.4%
  10. SWE-Bench Pro v2-Hard77.2%
  11. SWE-Bench Pro v165.5%
  12. SciCode49.7%
  13. DeepSWE v1.144.4%
  14. Humanity's Last Exam (no tools)36.2%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Strengths

  • Sparse MoE with 23B active of 501B total parameters
  • 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1 in Reflection AI's launch table
  • 97.8 on AIME 2026 and 90.5 on GPQA Diamond
  • Effective context length of 1M tokens
  • Weights announced under the permissive Apache 2.0 license
Reflection AI line chart of Beam's busiest-expert load relative to uniform routing over training progress, rising to nearly 2x early in training and falling to 1.04x by the end.
Beam's expert load balance during pretraining, averaged across MoE layers.Reflection AI ↗

Best for

  • Reach for it for agentic coding and terminal tasks once you have early access or the weights.
  • Reach for it for self-hosted coding agents once the Apache 2.0 weights are released.
  • Look elsewhere for image, audio or video input: Beam is a text-only model.

How to access

FAQ

What is Reflection Beam?

Beam is Reflection AI's first open-weight model, announced on October 5, 2026. It is a sparse Mixture-of-Experts language model with 501B total and 23B active parameters, built for coding, reasoning and agentic tasks.

Can I download the Beam weights?

Not at announcement. Reflection AI says it will release the weights under the Apache 2.0 license later in October 2026, together with a technical report, model card and developer artifacts. Until then Beam is open to a select group of users through a waitlist on platform.reflection.ai.

What is Beam's context window?

Reflection AI says midtraining extends Beam's effective context length to 1M tokens. Its reinforcement-learning stage used a maximum context of 256K tokens.

How does Beam score on coding benchmarks?

In Reflection AI's launch tables Beam scores 80.9 on SWE-bench Verified, 80.1 on Terminal Bench v2.1, 78.0 on SWE-bench Multilingual, 77.2 on SWE Bench Pro v2-Hard and 44.4 on DeepSWE v1.1. On Terminal Bench v2.1 the same table lists 81.0 for GLM 5.2 and 86.6 for Qwen 3.8 Max.

How was Beam trained?

Reflection AI pretrained Beam on 23.8 trillion tokens in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, then ran reinforcement learning that generated over 100 million rollouts on 10.5K GB300 GPUs over four weeks.