Research Radar 2026-10-07

Research Radar: Reflection's Beam and Mistral Large 4 Arrive a Day Apart -- Both Built on Asynchronous RL at Scale, Neither With Weights Yet

Reflection announced Beam, a 501B-parameter (23B active) text model, on October 5; Mistral opened a preview of Mistral Large 4 on October 6. Both describe the same recipe -- reinforcement learning with rollout generation and training decoupled -- and both promise open weights later this month. Until then, every benchmark is the vendor's own.

Two Western labs announced open-weight frontier models a day apart. On October 5, 2026, Brooklyn-based Reflection introduced Beam, "a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active." On October 6, Mistral opened a public preview API for Mistral Large 4 (ML4), which SiliconANGLE reports is a 1-trillion-parameter mixture-of-experts model using 49 billion parameters at a time. Neither has released weights: Reflection says weights, technical report and model card come "later this month"; Mistral says "by the end of the month."

The shared research trend: RL infrastructure is the product

  • Reflection says it "deployed 10.5K NVIDIA GB300 GPUs for four weeks generating more than 100 million rollouts" with up to 256K-token context, using about 1.3 billion sandboxes for training and grading and a million coding, agentic and STEM environments. Rollout generation and training ran asynchronously, at inference-to-training GPU ratios between 3.9:1 and 5.4:1, with new weights reaching the inference fleet in a median of about 12 seconds.
  • Mistral describes "an autoscaling fleet of actors" generating "tens of thousands of rollouts in parallel while model training proceeds asynchronously." At its current scale of about 3,000 GPUs, it says a run produces roughly 33 billion tokens a day, about 16 billion of them trainable. ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European data centres, and the RL run "is still in flight."

The common point: both labs present the environments, sandboxes and asynchronous pipelines -- not the base architecture -- as where their gains came from, and both say capability was still climbing when they shipped. Mistral says ML4 uses the same training and RL environment it offers customers through Mistral Forge; Reflection pitches "AI factories" where institutions train its models on their own data.

What each claims

Reflection says Beam matches Z.ai's GLM-5.2 on advanced reasoning while using "3–4× less inference compute," and concedes that "frontier open models like Kimi K3 remain ahead on raw capability." It also describes training a separate safety-and-alignment teacher model and merging the two by distillation, and promises to open-source its internal safety evaluations. Mistral reports 61.7% on DeepSWE v1.1 and a Coding Agent Index of 49.8%, ahead of DeepSeek V4 Pro and Qwen3.8 Max, and second of five in a blind Surge AI coding evaluation, behind Claude Opus 5. TechCrunch notes Reflection's claims "haven't been independently verified."

What remains uncertain. Licences and architecture details arrive with the weights (Reflection has said Apache 2.0; Mistral has not stated a licence on its launch page). Until then, "open-weight" is a promise, and no independent group can reproduce the numbers or test the safety training.

Reflection's Beam (October 5) and Mistral Large 4 (October 6) both credit large asynchronous RL pipelines -- 100M rollouts on 10.5K GB300s for Beam, 33B tokens a day for ML4 -- for their gains, but neither has shipped weights, so every benchmark so far is self-reported.