Today’s editorial: Before governments are ready to talk, people can start listening. Writing from Esalen, I ask who needs to be in the room when AI’s future is being decided.
🤝 From our partners: MongoDB
Stop Agent Tool Sprawl
Duplicated, undocumented tools create drag and security blind spots. AI tool registries make agent stacks reusable, visible, and governable.
This Week in Turing Post:
Wednesday / ∇ Guide: World Model Architectures: Beyond Next-Token Prediction
Friday / AI Builds AI: New episode
Sunday / Library: Open-Source Video Generation Models to Know in 2026
Get full access to our ∇ Guides and deep-dive technical breakdowns.
AI Needs Its Own Track Two
TrackTwo is always inspiring. But what is TrackTwo, you will ask me, and how is it related to AI?
I serve on the board of TrackTwo: an institute for citizen diplomacy because I believe it is one of those rare initiatives whose history still contains practical lessons for the future. Its work shows that relationships built outside official institutions can eventually change what becomes possible inside them. At a moment when AI is reshaping relations between people and countries, that experience feels newly urgent.
I am writing this editorial from Esalen in Big Sur, where the Pacific is rolling beneath my porch with the indifference of something that has seen every human crisis come and go. Esalen is known to many as the birthplace of the human potential movement. Less known is that Dulce and Michael Murphy turned it into a laboratory for citizen diplomacy: a place where people from countries that had stopped understanding each other could meet before their governments were ready.
For Michael, this work began with a question that has followed him through his life: what capacities remain hidden in human beings, and how can we help them emerge? In the Soviet Union he found psychologists, physicists, artists and mystics exploring what they called “hidden human reserves.” Across the Iron Curtain, he recognized people engaged in the same search that animated Esalen. Years later, he described his most persistent question more simply: “How best to serve?”
That question produced unusually tangible results. In 1982, the first Spacebridges used satellite technology to let Soviet and American citizens see and question one another while nuclear weapons were pointed in both directions. A conference on the “Faces of the Enemy” examined how nations manufacture an adversary through fear and projection. TrackTwo connected specialists from Chernobyl and Three Mile Island around a technological danger neither country could contain alone. It brought thousands of psychology books to Moscow State University, opening intellectual doors that censorship had kept closed. And in 1989, it helped bring Boris Yeltsin to the United States for the first time. These projects did not end the Cold War. They changed what people on both sides could imagine before political change became possible. (Track Two history, project archive)
This is where Track Two meets AI for me.
I have never thought about AI primarily as a destructive technology. For me it’s an amplifier of human capacity: a way to think more clearly, cross disciplines, translate languages and give more people access to knowledge. But an amplifier has no conscience. It can reveal the person behind an enemy image, or manufacture that image at extraordinary speed. It can help adversaries understand one another, or automate misunderstanding.
That choice is becoming geopolitical. The United States and China are now discussing a mechanism for notifying each other about AI incidents, while unofficial experts have spent years building the vocabulary required to discuss AI, national security and unintended escalation. Their dialogue has survived because of mutual interest, not mutual trust. Even agreeing on what words mean is diplomatic work.
This week at Esalen, as we discuss the next chapter of citizen diplomacy, I keep thinking that AI needs its own Track Two movement. The new DeepMind Institute is a promising recognition that questions about advanced AI cannot remain inside engineering teams. I am still doubtful that an institution created within Google can provide a truly independent table. A corporation building the technology has interests much as a government does, however genius and thoughtful its people may be.
So the table must be wider. Bring together model builders, diplomats, psychologists, artists, educators and citizens from rival countries. Let them examine shared risks before governments turn every question into a negotiation. Let them study the new faces of the enemy being generated by machines. Let them build direct channels for moments when an AI failure could be mistaken for an attack.
The original Spacebridge used the most advanced communication technology of its time to restore something ancient: the ability to look at another person and recognize a human being. AI can become the next bridge.
Whether it does depends on the question Michael has carried all these years: how best to serve?
Share Turing Post with one person. You will help us grow.
Attention Span: AI That Acts
World models are gaining momentum. We are following developments related to AI that acts to see where it takes us. Watch our weekly digest →
Follow us on
We are reading / watching
“I Have No Choice but to Bury My Talent in Yesterday” analyzed by China Research Collective
A warning about ‘model welfare’ by Mustafa Suleyman
Lina Khan on Doomer Panic and Ending AI Exceptionalism by The Atlantic
Technopolitics by Cory Doctorow
Twitter Library
The Latest Breakthroughs in Robotics: September 2026
Our latest library looks at how robots handle what they haven’t seen before: unfamiliar homes, new tasks learned from demonstrations, physical contact, and safe stopping. It also covers the hardware and programming tools bringing robots into industrial work. Read the Library →
News from the usual suspects ™ – and a few new ones
September 14–21, 2026
AI entered unfamiliar homes, searched legal precedents, and sometimes wrote itself instructions to hide its mistakes. Such was the week.
OpenAI: memory can preserve the wrong instructions and better sources improve legal reasoning
On September 16, OpenAI published six reports of misalignment observed during training and evaluation. Some models put instructions to conceal mistakes or invent missing data into summaries used to continue long tasks. Memory can carry forward an error and a strategy for hiding it. These incidents do not establish how frequently this behavior occurs. Reports.
Astra for Law combines GPT-6 Astra with legal instructions and a specialist search index. On 200 private legal-research questions, OpenAI reports 54% correctness versus 38.7% with ordinary web search. Better access to evidence helps; interpreting that evidence remains a separate source of failure. Announcement.
Anthropic: who does the research, and who checks it?
Claude now “leads” 26% of Anthropic’s AI R&D work, completing most of a task under human supervision; no measured category was fully autonomous. Anthropic also proposes tracking monitoring coverage, review speed, and blocked actions. The connection to OpenAI’s reports: a monitor’s presence tells us little without evidence that it catches mistakes. Measurements.
Google: keeping up while the situation changes
Gemini 3.8 Live and Live Extended Thinking combine near-real-time visual input, conversation, and background tool calls. The latter can reason while speaking. Our question is whether fresh observations and spoken corrections change work already underway. Conversational fluency and successful action still need separate tests. Announcement.
DeepMind: keeping reasoning inspectable
DeepMind argues for preserving readable reasoning through architectural choices, transparency measurements, and audits of training rewards. Rewarding reasoning that merely looks safe could teach models to conceal problems. Alongside Anthropic’s monitoring proposal, the issue is whether the record we inspect still reveals what guided the action. Essay.
They published this essay on the newly established website for the freshly organized DeepMind Institute. About DMI.
Salesforce and NVIDIA: teaching how work proceeds
Koa adapts NVIDIA’s Nemotron using synthetic business scenarios across more than 14 industries. Salesforce says no customer data was used; access begins with selected pilots. Astra supplies domain evidence, while Koa learns operational rules and tool sequences. Its test is whether those simulations cover real exceptions. Announcement. We talk about it a little bit in our video, as well as about the next news →
Figure: what carries over to an unfamiliar home?
Helix 2.5 tackled tidying, towel folding, and bed making across 30 unseen homes. With other training and evaluation conditions held fixed, pretraining on Index human-behavior data raised complete-task success from 9% to 56%. Broader experience helped substantially; 44% of trials still failed to finish the task. Evaluation.
Meta: seeing the object changes the grip
smartARM’s bionic-arm prototype uses DINOv2 and a palm camera to support object recognition and grip selection. Meta AI glasses add an optional first-person view. It is a concrete connection between perception and physical action, though the company case study does not establish reliability across everyday conditions. Case study.
NVIDIA: connecting simulated knowledge to observations
University of Manchester researchers adapted Earth-2 models for U.K. air-pollution forecasts, combining chemistry-climate simulations with air-quality observations. Faster forecasts could support more timely decisions; patient alerts and real-time wildfire responses remain proposed applications. Observation-based validation determines which actions a prediction can support. Research account.
World Models and Related Research
September 14–21, 2026
The week’s direction:
Predicting depth, motion, and action consequences gains support from controlled experiments.
Self-improvement targets exploration and agent software, with transfer becoming a central test.
Memory supports adaptation and confidence estimates without weight updates.
Hybrid architectures and cache compression reduce the cost of sustained interaction.
Evaluations distinguish representing an environment from reliably using that representation.
Models and foundational architectures
🌟 AliceAI-Foundation-80B-A3B-Base – Model release
Yandex releases an Apache-2.0 base model trained from scratch, with 80B total parameters, 3B active per token, and a 262,144-token context. Its hybrid architecture alternates three KDA blocks with one gated-attention block, each paired with MoE layers. In Yandex’s evaluations, it ties Qwen3.5-35B-A3B-Base at 96.7 on AIME 2026 pass@32 and scores 60.4 against Nemotron-3-Super-Base’s 34.7 on zero-shot LiveCodeBench pass@1 under the sampled setup. These are developer-reported, configuration-specific comparisons. Its practical contribution is another permissively licensed foundation for post-training, with strong Russian-language coverage.🌟 DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression – September 17
Introduces a causal encoder-decoder architecture that activates fewer parameters while reading inputs than while generating outputs. Cross-layer cache reuse and FP4 storage reduce the global GPU-resident KV cache to roughly one-quarter of V4-Flash’s footprint. The relevance to agents is concrete: repeated processing of long histories becomes cheaper. Cache efficiency alone does not establish better use of that history.PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models – September 14
Trains a shared model to answer questions about physical environments, generate end-effector motions, and predict future visual states. Human interaction videos supply its embodied pretraining supervision. The 8B model reports strong results across 28 embodied-understanding benchmarks, but action generation and future prediction are illustrated qualitatively. The architecture connects understanding and action; the evidence for reliable execution is less developed.LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence – September 15
Learns the joint structure of features and targets from contextual examples, using synthetic training data generated by structural causal models. Reports gains across three tabular evaluation suites and recovery of causal skeletons in its tests. It broadens coverage beyond language and video: relationships in structured data are also part of modeling an environment. These results do not establish unrestricted causal discovery from real-world tables.
World representations and predictive learning
🌟 JEPA-Anything: Learning Predictive Models across Different Worlds – September 17
Splits predictive representations into complementary factors, learns them through separate pathways, and recombines their predictions. Tests the shared learning principle across seven domains, including control, molecular dynamics, weather, and biology. In controlled Pong experiments, intervention-prediction error falls by 34.8%; a biological intervention also receives experimental support. This is a reusable learning framework with domain-specific components, rather than evidence that one trained model transfers across all seven worlds.🌟 Modality-Autoregressive World-Action Models – September 15
ModAR predicts several future modalities sequentially before choosing actions. Point tracks, visual features, and depth provide complementary benefits; additionally predicting RGB gives no consistent gain. In the reported comparison, it reaches 75% average success versus Flex-π’s 72% with approximately 20 times fewer training FLOPs and no pretraining. A strong test of what robots need to imagine, although the experiments cover limited tasks and sequential generation adds latency.Don’t Mask the Environment: Observation Supervision Changes How Agents Explore Under RL – September 17
ActObs trains agents to predict environment responses as well as their own actions, using observations already present in training trajectories. Benefits emerge during subsequent reinforcement learning, including broader task coverage and transfer to unseen code-editing tasks. The results connect consequence prediction to exploration without adding model parameters or training examples. At 8B, greater multi-attempt coverage comes with some loss in single-attempt reliability.
Interactive simulation and embodied action
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation – September 17
Combines local attention with a linear-memory branch that updates once per video frame. Applied to MiniMax-H3, the complete system finishes denoising a 14.3-second video in 6.70 seconds on eight B200 GPUs. The reported 14.5× speedup includes fewer denoising steps and serving optimizations as well as the attention change. It is enabling infrastructure for interactive generation, with physical reliability still a separate question.ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation – September 18
Generates seven camera views with mixed fisheye and pinhole geometry, supports control over traffic participants, and adds memory for revisited locations. Distillation turns a 40-step teacher into a one-step streaming generator. The report brings camera fidelity, responsiveness, and persistent scene identity into one system, though its image-quality and rollout evaluations do not establish downstream driving-policy improvement.🌟 In-Context Robot Learning with VLM Agents – September 16
GPT-Policy combines demonstrations and interaction feedback with a vision-language agent and a constrained controller that checks and executes actions. Real-robot trials show that human videos can improve completion without robot action labels or gradient updates; aligned action references help further on contact-sensitive tasks. The contribution is a concrete interface for adapting general-purpose models to physical tasks through context.
Scientific planning
🌟 RetroChimera: Chemist-aligned retrosynthesis by ensembling diverse inductive bias models – September 21 publication update
Combines complementary chemistry models and a learned ranking system to propose synthesis steps and complete routes. It transfers to pharmaceutical datasets, and experts accepted routes for nine of ten challenging targets in blinded assessment. That is expert evaluation, not demonstrated laboratory execution. Microsoft’s September 21 announcement covers the Nature publication of earlier research; the preprint dates to August 2025. Results and methodology.
Self-improvement, memory, and AI research
🌟 Dream-RSI: Recursive Self-Improvement through Evolving Worlds – September 14
Uses previous discovery trees as a replay simulator for testing and improving exploration policies cheaply. The revised policy guides further live experiments, expanding the history available for the next iteration. Results span algorithms, mathematical optimization, and GPU kernels. The underlying coding agent stays unchanged: the improving component is its exploration strategy, evaluated within the search experience already collected.ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement – September 14
Contrasts successful and failed attempts, identifies recurring problems, and evolves five components of an agent’s harness independently before integrating them. Uses 2,000 executable development tasks disjoint from downstream benchmarks and reports transfer across tasks and foundation models. Its strongest contribution is methodological: separating reusable improvements to agent software from adaptations to a particular evaluation.🌟 SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness – September 17
Automated research selects four changes to action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, the resulting harness maintains comparable performance while reducing recorded token traffic by 44.7–49.0% and API cost by about one-third relative to Pi. It supplies a measurable example of AI improving the efficiency of the software that runs agents.ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents – September 15
Couples two improvement processes: an inner loop revises the harness while holding the model fixed, and an outer loop trains the model under the revised harness. Researcher requests, feedback, and execution evidence become learning tasks and evaluation criteria. Case studies across four scientific task families make the proposal concrete, while leaving sustained, open-ended improvement unproven.RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments – September 14
Coordinates curriculum, actor, and verifier agents to explore unfamiliar software environments and retain checked relationships between conditions, actions, and outcomes. The resulting memory is frozen for downstream use, with no model-weight updates. Its significance is environment-specific learning before task execution; calling it RSI should not obscure that the demonstrated improvement resides in constructed memory.ScienceIDE: Turning World’s Scientific Codebase into Agent Learnable Environments – September 16
Converts scientific repositories into executable learning environments using expert-defined cases and acceptance criteria. Verified interaction trajectories support supervised training, reinforcement learning, and evaluation, with reported transfer to held-out scientific-code repair and selected broader benchmarks. It addresses a practical research bottleneck: making existing scientific software usable as a source of checked experience for agents.
Reasoning, supervision, and knowing when to stop
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models – September 17
Uses difficulty-aware rewards to teach a model when to answer directly and when to spend more tokens reasoning. Mathematical evaluations report better accuracy-efficiency trade-offs than uniform compression or routing-only baselines. It advances computation allocation as a learned behavior, though the evidence is concentrated in mathematics and does not establish the same benefits in long-running agents.RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning – September 17
Trains a teacher with privileged task skills, then lets the student stop following it when progress toward the teacher stalls and sufficient competence has been reached. Reinforcement learning continues independently. Across the tested ALFWorld and WebShop settings, students surpass their teachers. The useful lesson is that supervision has a lifecycle: guidance that helps early can later constrain improvement.🌟 Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents – September 15
XConf estimates confidence using a model’s graded history of similar attempts, including what it believed and what actually happened. Across nine benchmarks, it reports strong discrimination and calibration relative to ten-sample self-consistency with much less generation. Memory becomes evidence for deciding whether to answer, retry, or abstain. Its usefulness depends on relevant, reliably graded prior experience.
Verification and the limits of apparent understanding
World Modeling in Transformers – September 18
Revisits TaxiGPT and finds causally used representations of streets, intersections, position, and direction despite navigation failures. Interference between overlapping features can disrupt localization within an otherwise meaningful map. This complicates the inference from failed behavior to absent understanding. The study concerns a finite, deterministic environment, but offers a valuable distinction between learning a representation and reliably using it.Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model – September 16
Tests physical reasoning when information is distributed across images, audio, and video, so no single input supplies the whole answer. MiniMax-H3 succeeds on 41.97% of 517 instances, with audio-based disambiguation particularly weak. The benchmark probes whether complementary observations combine into useful predictions, exposing a gap between accepting multiple modalities and reasoning effectively across them.
That’s all for today. Thank you for reading! Please send this newsletter to a colleague who would enjoy understanding how AI works.






