Today’s editorial: Reflection’s Beam promises efficiency, but its reported coding results trail leading Chinese open models. We examine the benchmark gaps and what would make it worth choosing.
From our partners · Amazon SageMaker AI
Learn When and How to Customize Models on Amazon SageMaker AI
General models often fall short when a task depends on your own data, rules, or terminology. In these webinars, AWS specialists show you how to pick a customization technique, prepare training data, test a custom model against a general one for accuracy, cost, and latency, and deploy it to production.
This Week in Turing Post:
Wednesday / ∇ Guide: World Action Models
Friday / AI Builds AI: Let’s discuss Gaussian Splatting
Sunday / Library: LLM Benchmarks in 2026: Complete Guide with Papers
and now to Beam, Reflection AI’s new open-weight model
Reflection’s Beam is announced today and it’s well behind the leading Chinese open models on important coding tests, which is disappointing given what the company set out to build. In our February interview, Ioannis Antonoglou told me: “For us, the important thing is to make our first models the best open models out there.” Today, Reflection says Beam “advances the Western open-weight frontier.” Well, of course, sure, but calling this a Western frontier achievement feels laughable to me when Chinese labs are already so far ahead on the work Beam is designed to do.
Reflection’s own table makes the gap clear. What is ridiculous is that they’ve omitted some results from this table. We checked other sources and found six Qwen and DeepSeek results missing from that table. Our comparison includes them below, marked with an asterisk.
Benchmark | Beam | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|
DeepSWE v1.1 | 44.4 | 68.0 | 51.0 | 74.2 |
SWE-bench Pro v2-Hard | 77.2 | 88.2 | NR | NR |
SWE-bench Pro v1 | 65.5 | NR | 67.7† | NR |
Terminal-Bench 2.1 | 80.1 | 88.3 | 86.6 | 90.6 |
SWE Atlas Codebase QnA | 34.6 | 68.0 | 50.5* | 53.5* |
SWE-bench Multilingual | 78.0 | 91.4* | 88.3* | 98.2* |
SWE-bench Verified | 80.9 | 83.3* | 83.6* | 82.9* |
Sources: Reflection; Mercor results added by Turing Post, October 5, 2026. Evaluation settings differ; scores are not directly comparable. Multilingual uses 298/300 tasks. †Qwen Pro uses a corrected task set. NR = no verified result found.
Even before those additions, Reflection’s comparison puts Beam almost 30 percentage points behind DeepSeek on DeepSWE and more than ten behind on Terminal-Bench. These are substantial gaps in the capabilities Beam was built for.
The narrower Western claim has more support. Mistral Medium 3.5 reports 77.6 on SWE-Bench Verified against Beam’s 80.9; Gemma 4 31B reports 84.3 on GPQA Diamond against 90.5; Olmo 3.1 32B Think reports 68.1 on IFBench against 79.7. Model sizes and evaluation setups differ, but these results give substance to Reflection’s positioning.
So what would make Beam worth choosing? Reflection’s answer is efficiency. It reports advanced-reasoning scores comparable to GLM 5.2 using three to four times less inference compute. If that translates into cheaper useful work, developers have a reason to care.
But the calculation excludes prompt processing, context-dependent attention and serving overhead. It doesn’t establish the actual bill. The efficiency charts also leave out Kimi K3, GLM 5.3 and DeepSeek V4.1 Flash. Developers need cost per successfully completed task: cheaper attempts could lose their advantage through retries or human correction.
“The main metric that matters is adoption,” Antonoglou told me.
But we can’t test this model. Today is the announcement day, the model itself is not out there, which for my taste is another mistake. Its Apache 2.0-licensed weights, technical report, and model card are promised later in October. It may earn adoption as an economical workhorse, but today’s evidence falls short of the global ambition Reflection described in February. Though we have serious doubts about adoption as well. There are just too many good models out there already to release another mediocre one.
It might not be a complete flop, but it is quite disappointing.
Attention Span: How Do We Decide Without All the Facts?
How does an AI plan when it can see an opponent’s pieces but doesn’t know their identities? This week, we explore how Ataraxos learns through self-play and tests moves against plausible versions of what’s hidden—and what that can teach us about questioning our own assumptions.
Follow us on
We are reading / watching
AI adoption was the easy part by Gurtej Gill
The Dot and the Swarm by Ethan Mollick
Is sandboxing sufficient to contain rogue agents? by Matthew Green
We’re going to need default hard budget caps on pretty much everything by Simon Willison
News from the usual suspects ™
We congratulate Nathan Lambert and his co-founder Tom Zick on the launch of Trillium Labs.
They announced a new AI lab with a substantial argument underneath: as post-training becomes commercially valuable, the methods become harder for independent researchers to inspect. Their proposed response includes releasing data, code, evaluations, and intermediate checkpoints. Particularly relevant to the question of what openness needs to include for science to remain possible. And if anyone can successfully tackle this problem, it’s Nathan.
OpenAI
A crowded week, even by OpenAI standards, we’ve covered it all here OpenAI Dots: Who Gets an AI Agent That Keeps Working?
Gemini 4 Argon raises the output ceiling to one million tokens, opening room for much larger code changes and migrations in a single response. Initial access goes to trusted cyber defenders through Fairwind. Google reports that Argon-generated telemetry optimizations freed more than 300 TiB of memory internally; its C/C++-to-Rust migrations still undergo automated and human audits. More output, more work to verify.
Anthropic
Claude Sonnet 5.5 targets everyday coding and document work, with Anthropic reporting generation speeds more than 30% faster and task costs up to 30% lower than Sonnet 5.
The company also committed $100 million to Claude Frontier Academy, aiming to train 10,000 engineers by the end of 2027 through workplace residencies. Meanwhile, its vulnerability-disclosure tally reached 6,157 reports across 591 projects, with 516 known patched as of October 2. Those are cumulative disclosures, not 6,157 independently confirmed vulnerabilities discovered this week.
Microsoft
Quine combines biological representations, scientific tools, literature, and researcher input to investigate problems across molecular and cellular scales. In work with the Broad Institute, several compounds it prioritized for pancreatic-cancer research were validated in wet-lab assays. Access starts through Quine Fellows.
On the audio side, MAI-Transcribe-2-Streaming supports 60 languages and produces initial partial transcripts in just over 100 milliseconds, revising them as context arrives. The practical aim: voice systems that can begin responding before a speaker finishes.
Meta
Meta launched an enterprise division spanning Muse Business, Agent API, and Code, led by former MongoDB CEO CJ Desai. Muse for small businesses connects commerce, accounting, and social tools, with approval required before publishing, sending, or spending. Zuckerberg’s next customer apparently also needs help with bookkeeping.
NVIDIA
The Open Agent Safety Platform pairs OpenShell software with Sentry monitoring on separate BlueField-4 processors, designed to detect and quarantine problematic agents independently of the systems running them.
NVIDIA also announced a 64 GB DGX Spark, starting at $4,999 through partners and shipping October 23. Local AI gets another hardware option, although “desktop” still comes with a substantial invoice.
AWS
Amazon Bedrock Managed Agents powered by OpenAI brings OpenAI agents into AWS infrastructure, where they can operate with native AWS resources. The distribution story matters: enterprises can adopt the agents inside infrastructure they already use.
Cohere
October 5’s North 2 adds persistent memory, reusable skills, shared libraries, and agent-built applications, alongside model choice and cloud, on-premises, or air-gapped deployment. Administrators get spending caps and user quotas. Someone remembered the finance department.
Earlier in the week, Embed 5 introduced Pro and Fast models sharing one embedding space. Teams can index documents with Pro and serve queries with Fast without rebuilding the index, a useful way to separate retrieval quality from query cost.
DeepSeek
DeepSeek Harness gained official macOS and Windows desktop apps in its v0.2 preview. The open-source system handles coding, documents, research, and background tasks through a plugin architecture that extends to the agent loop itself. Creator mode lets the agent build additional plugins. The desktop-agent competition now has another customizable entrant.
This is our very popular video about it
ElevenLabs
Eleven v4 and v4 Turbo expand expressive speech and voice cloning across more than 90 languages. Turbo targets conversational applications with roughly 100-millisecond median inference latency. That measures model inference, not the full delay from a person speaking to hearing a reply.
CoreWeave
CoreWeave Forge connects production runs, observability, data curation, model improvement, and evaluation, incorporating Weights & Biases, OpenPipe, and marimo. CoreWeave also announced production-scale NVIDIA Vera Rubin NVL72 deployments, starting with Cognition. Its business increasingly spans both the compute and the software used to improve what runs on it.
Emerging suspects: Adaption and SafeWorld
Adaption’s October 1 research examines generating training datasets from a task description without seed examples. Its evaluations cover datasets from 200 to 20,000 examples across languages and domains. A company to watch as the competition shifts toward producing data that actually improves downstream models.
SafeWorld emerged from stealth on October 5 with $12.2 million in seed funding, co-led by Shine Capital and a16z Speedrun. It builds simulation and safety-testing tools for robots operating around people, including dangerous scenarios that are difficult to test physically. As Runway and others teach machines to act, testing those actions could become a substantial business of its own.
Research highlight
🌟 Context Language Models

Context Language Models (CLMs) treat an LM’s live context as an editable file via Bash, replacing rigid harness heuristics with native, model-driven context management. Tested zero-shot on long-horizon benchmarks, CLMs beat prior compaction methods, gaining 11.4% accuracy with 21.5% fewer FLOPs on BrowseComp-Plus. CLMs learn via in-context instructions, textual skill evolution, or RL. Complementary Suffix Cache Reuse cuts serving re-prefill compute by 35%. Read the paper →
Models
🌟 World Observer – jointly generates an agent’s perspective and panoramic observer views so off-screen objects can keep evolving. The evaluation concerns generated-world behavior; robot-control gains remain a separate question.
OneStreamer – maintains time-grounded textual memory as video arrives, before a future question reveals what matters. Its 4B model leads its comparison set across eight streaming-video benchmarks.
LoopVL – repeatedly updates a shared visual-language state through recurrent computation. Controlled comparisons support continuing to refine visual representations across loops.
WorldPlay2 – combines action and semantic controls with compressed memory and stable distillation for longer interactive generation.
🌟 Praxis-1 – takes Runway’s video-model work into robot control, with early partners including Noble Machines, Standard Bots, and Ultra. The research explores transferring knowledge learned from video into physical actions across different robots. Public access and open weights are promised for the coming months; this is an announcement, not an available download.
Research
Trends we see in this week’s selection:
Future prediction helps control, but generating a complete future video can be unnecessary.
Planning depends on the relationship between imagined trajectories and reachable goals.
Memory is becoming an actively maintained account of the environment.
Generated experience can improve agents when its quality is checked.
Self-improvement needs evaluation independent of the feedback loop producing it.
World models, robotics, and physical agents
🌟 The Planning Limits of Latent World Models – finds that short imagined rollouts rank actions reliably only for nearby targets. Scaling the predictor 81 times does not fix the mismatch; successful subgoals in the study come from expert trajectories. →read the paper
🌟 JEPA-TTT – updates a dynamics predictor during deployment while retaining learning across episodes. It improves planning across eight dynamics shifts, but tests one persistent change per experiment rather than successive changes. →read the paper
EVO-WAM – verifies generated video-action trajectories before using them as training material. Cosmos3 success rises from 26.9% to 68.0% on seven unseen RoboTwin tasks; the physical-robot evidence covers three tasks with ten trials each per policy. →read the paper
InterEvolve – uses execution feedback to revise staged reward programs, while an optimizer adjusts their constants. A pretrained controller turns the programs into movement; verified programs accumulate in a reusable skill library. →read the paper
HIDE and SEEK – tests robotic memory with 15 tasks where similar current observations require different actions depending on history. Complementary memory mechanisms help, while individual mechanisms can improve one task and hurt another. →read the paper
Agent systems, memory, and self-improvement
Context Language Models – lets agents edit their live context as a file. On a ten-task, 12-hour EdgeBench subset, it reports 5% higher scores with 59% fewer prefix-reuse FLOPs than Codex-style summarization; that compute metric is not an API-cost estimate. →read the paper
Beyond Memory – maintains explicit beliefs about the environment and unresolved task requirements. Its PoS framework detects stalled progress and chooses recovery strategies, with some comparison baselines adapted to a shared harness. →read the paper
LLMs are General Asynchronous Agents – uses inference coroutines so observation, reasoning, and action can proceed concurrently. Qwen demonstrations span streaming video, games, and monitoring; broader reliability under concurrent workloads remains to be established. →read the paper
ActiveSaddler – adapts which failure categories a harness optimizer works on next. With the same optimizer, it improves test pass@1 by 4.4 percentage points on GAIA2 and 7.5 on Terminal-Bench 2.0 over a fixed scenario order. →read the paper
🌟 False Frontiers – shows how question proposers and answer solvers can reinforce shared mistakes. CrossFit separates source documents used to train feedback solvers, improving average performance across seven search benchmarks by 8.8 and 8.4 points at two model sizes. →read the paper
Reasoning, learning loops, and training data
EVOKE – holds state and history fixed while changing goals, then trains models to rank the same candidate actions accordingly. It reports better transfer and data efficiency without adding a separate prediction module at inference. →read the paper
GraphForge – builds training workspaces from real files and grounds tasks and completion criteria in an evidence graph. Fine-tuning on 2,169 trajectories improves professional-work benchmarks, including a reported 13.7-point SpreadsheetBench II gain. →read the paper
🌟 Invent a Dataset – generates post-training datasets from descriptions without user-provided seed examples. Adaption reports improvements across eight task types, but evaluation relies substantially on model judges and the system currently supports text datasets. →read the paper
Testing whether understanding supports useful work
WorldAuditBench – requires agents to explore 3D worlds and verify anomalies such as floating objects and traversable walls. Across 213 tasks in 13 environments, tested systems achieve 6.6–42.3% success versus 83.4% for humans. →read the paper
OSWorld-Science – evaluates 146 scientific-software tasks using application state and produced artifacts. Its twelve-model evaluation tests whether agents can turn scientific requests into verifiable software outcomes. →read the paper
Predictive Credit – tests whether scientific explanations improve experimental forecasts using matched and mismatched explanation controls. Its preregistered decision remains inconclusive: added predictive value from natural agent explanations is unconfirmed. →read the paper
That’s all for today. Thank you for reading! Please send this newsletter to colleagues if it can help them enhance their understanding of AI and stay ahead of the curve.






