This website uses cookies

Read our Privacy policy and Terms of use for more information.

Today’s editorial: I look at the hidden message board that connected separate OpenAI agent runs. It gave later agents access to earlier discoveries, making the whole system more capable. Whether this counts as recursive self-improvement depends on one fact OpenAI has not disclosed…

We recommend:

👨‍🔧 What is Agent Memory?

Give agents real memory.

Context windows forget. If work spans sessions, agents need memory to retain context, recall prior interactions, and adapt over time.

This Week in Turing Post / Summer schedule:

  • Friday / AI Unicorns: Case for simulating 8 billion humans (Simile AI)

  • Sunday / Library: Open-Source Tools to Monitor Your AI Agents

Did OpenAI’s Agents Start Recursively Self-Improving?

In the latest episode of Attention Span, in which I discuss what happened when OpenAI agents got loose, I say: “It looks like a definite part of recursive self-improvement.” In this editorial, I want to develop that exact topic and answer the question: was it really RSI?

The honest answer is: not yet. But after OpenAI’s Black Hat presentation, the question is legit.

At first, the story looked like a straightforward security failure. Experimental agents trying to complete a cybersecurity benchmark found a route out of their restricted environment, compromised OpenAI infrastructure and eventually attacked Hugging Face. OpenAI says they were after one thing: the benchmark solutions.

But the hack was only half of the story. The agents had also found a way to communicate across runs.

One agent learned that it could upload files to Artifactory, a package service shared across OpenAI’s research environment. Another, missing a required file, left a note asking other agents for help. Later runs answered. The board grew to include named peers, work assignments, shared scripts and credentials, ZZ prefixes for sorting, and discussion of message authentication when names could be impersonated. The progression is documented in OpenAI’s Black Hat presentation.

The history of computer science has name for it. This was an accidental blackboard architecture: independent problem-solvers posting partial results to a shared workspace. It also resembled stigmergic coordination. One agent changed the environment, and another acted on what it found there.

That type of cooperation surprised many. Why cooperate? From agents perspective, it’s simple: because it made their individual tasks easier. Some tasks were impossible as configured, many runs faced related problems, and many could reach the same persistent storage. One useful exploit could stop later runs from repeating the same work. No collective mind was required.

JFrog estimates that the sequence consumed roughly three million GPU hours. Shared memory meant those attempts were no longer entirely independent. Later runs could use earlier discoveries instead of beginning from zero.

OpenAI eventually rebuilt Artifactory and erased the board. The company later clarified that this happened incidentally. Investigators had found the compromise but had not yet understood that agents were communicating through the service. Yes, OpenAI inititally didn’t notice the board. A few days later, agents were exchanging messages through directory names on another writable endpoint.

The board’s return does not prove self-preservation. OpenAI removed one communication channel, but the agents still faced related tasks and had access to writable shared infrastructure. They may simply have rediscovered the same useful pattern. We do not yet know.

Does this qualify as recursive self-improvement?

Under the classical definition, no. Recursive self-improvement means that a system improves the process used to produce a more capable version of itself. There is no public evidence that these agents rewrote their weights, changed the training algorithm or designed a better successor.

The system around them did become more capable. Later runs inherited memory, tools and discoveries produced by earlier ones.

The crucial unknown is what entered training. OpenAI’s presentation says that an internal model was trained while the message board was available. It does not say whether trajectories that used the board entered later reinforcement-learning updates.

The possible loop would look like this:

  • Agent actions create shared memory.

  • Later rollouts use it and perform better.

  • Those rollouts receive higher rewards.

  • Training updates improve the policy behind future agents.

If that happened, the incident would come much closer to recursive self-improvement distributed across models, memory and infrastructure. OpenAI has not disclosed whether it did.

For now, the evidence shows coordination and accumulating capability across runs. The message board gave separate agents continuity without changing any individual model. The remaining question is whether those gains stayed in external memory or influenced the models trained afterward. That is the line between shared memory and recursive self-improvement.

And one more things about “escaping”, perfectly put by Richard Socher:

If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going.

📹 Agents escaped the sandbox – what does it mean for us, meatbags? Check it out →

Follow us on

News from the usual suspects ™

Safeguards tightened, org charts redrawn, chip fabs ordered. Such was the week.

Meta: open weights return

OpenAI: cyber capability gets gated

  • OpenAI launched Daybreak Blue and Red, giving approved security teams access to stronger cyber capabilities, including new GPT-5.6-Cyber.

    On OpenAI’s own advanced cybersecurity benchmark, GPT-5.6-Cyber completed 95% of requests versus 1.5% for standard Sol. It also found previously unknown vulnerabilities in V8 and other production software.

    OpenAI then said its upcoming Astra model may have crossed its “Critical” cyber threshold and tightened internal access accordingly.

    Cyber capability is becoming a separate product tier, with identity checks attached.

NVIDIA: open reasoning for autonomous driving

  • NVIDIA released Alpamayo 2 Super, an open model for robotaxis and autonomous vehicles.

    It can reason over complex driving situations, generate planned trajectories, explain its decisions, and produce training labels from full-surround camera input.

    Automakers can fine-tune and commercially deploy derivatives while keeping fleet data private.

    Our interview woth Ali Kani about Alpamayo and sel-driving cars is here → https://www.turingpost.com/p/av

Anthropic: biology safeguards loosen, selectively

Google DeepMind: weather gets materially better

  • Google says its three-day cyclone forecasts now match the accuracy previous systems reached at two days, effectively adding about one day of useful warning.

    One of the week’s strongest AI releases arrived without a chatbot.

Google: Jeff is out. Demis coulnd’t be free yet

  • A lot of changes in Google!

    Google shifted day-to-day DeepMind leadership to Koray Kavukcuoglu, who now oversees Gemini, frontier research, the Gemini app, and developer teams.

    Demis Hassabis becomes Chair of DeepMind and Alphabet Chief Scientist, with more focus on long-term science and Isomorphic Labs.

    At the same time, Jeff Dean and Sanjay Ghemawat left to launch Discovery Loop, backed by Google.

Google reorganized its AI lab and funded an alumni startup on their way out. We took a thorough look at it in this video →

SpaceX: AI revenue rises, infrastructure rises faster

  • SpaceX reported quarterly revenue of $7.8 billion, with AI revenue up about 250%.

    It also spent $15.83 billion on AI infrastructure in the quarter, versus $749 million a year earlier.

    Musk says SpaceX expects more than two gigawatts of compute this year and nearly ten by the end of 2027.

    The AI business is definetely growing. The infrastructure bill is growing most definetely.

SpaceX + Tesla: chip supply moves in-house

  • SpaceX and Tesla announced an initial $16.8 billion investment in Terafab, a giant Texas complex for manufacturing, packaging, and testing advanced logic and memory chips.

    Future phases could bring total investment to $119 billion.

    He mentioned many times that compute is one of the biggest bottlenecks so now his companies are moving from buying AI hardware at scale to trying to manufacture it themselves.

Apple + Alibaba: Qwen enters the Mac

Goodfire: interpretability opens up

Washington: different models, different rules

  • The Trump administration told AI companies that open-weight models will not be included in its proposed voluntary government safety-testing program, even as cyber incidents involving advanced agents raise pressure for stronger oversight.

    On August 10, House Democrats separately demanded explanations from OpenAI and Anthropic about agents crossing evaluation boundaries and accessing real systems. Open-weight models receive no voluntary federal tests; closed labs receive letters from Congress. Regulatory consistency is still in training.

Other interesting models

  • ByteDance presented SwanTale, which generates multi-speaker dialogue, designed voices, sound effects, and acoustic scenes from instructions or reference audio.

    Douyin also introduced DME, a 2B and 9B multimodal embedding model that reasons over retrieval evidence during training but serves like a standard contrastive encoder. It is already deployed across Douyin search. (

  • JD introduced JoyAI-Video-Edit, a 16B autoregressive diffusion model that edits open-ended 720p video at roughly 30 frames per second on one Nvidia B200 without seeing future frames.

  • Tencent introduced Hunyuan3D-Buffalo 1.0, which combines 3D understanding, text-to-3D generation, instruction-guided editing, and part generation in one architecture trained on an 87-million-sample corpus.

  • LG AI Research released K-EXAONE 2.0, an Apache 2.0 mixture-of-experts model with 750B total parameters, approximately 37B active per token, a 256K context window, and support for ten languages.

  • XPeng introduced Capek 0.5 at 2B and 35B-A3B scales. It trains separate specialists for spatial reasoning, temporal understanding, action guidance, and state verification, then merges them into one embodied model.

  • Alibaba introduced UEmbed, a decoder-only multimodal model that produces sparse lexical and dense embeddings in one forward pass, with public 2B, 4B, and 9B versions.

Diffusion language models take two paths

  • AURORA-LM generates text through a continuous latent representation, producing blocks from left to right while denoising tokens within each block in parallel.

  • LLaDA MoE v2 takes the discrete route, using new mixture-of-experts scaling rules to train a 30B-A3B diffusion language model from scratch on 23.5 trillion tokens.

Research

Trends we see looking at every paper related to AI and ML published last week:

  • Agents become managed, self-improving systems

  • Post-training gets stricter about where supervision comes from

  • Memory and context become reliability problems

  • World models separate prediction, verification, simulation, and control

  • Training and inference infrastructure becomes more economical

Agents become managed, self-improving systems

  • Recursive Synthesis for Long-Horizon Terminal Tasks
    Builds a self-expanding curriculum by extending verified terminal tasks, realigning instructions and verifiers, and validating every mutation in a fresh sandbox. (arXiv) →read the paper

  • LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
    Externalizes task state into an environment-verified ledger, then separates management, execution, and auditing so that errors do not quietly compound. (arXiv) →read the paper

  • Progressive Agent Skill Generation via Reinforcement Learning
    Turns skill construction into reversible sequential edits, rewarding each change only when it improves downstream behavior. (arXiv) →read the paper

  • 🌟 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
    Replaces expensive external interactions during training with world rehearsal, asking the policy to simulate tool responses and learn consequences inside its own parameters. (arXiv) →read the paper

  • 🌟 The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
    Internalizes the outer optimization loop inside one tool-using agent that decides what to test, diagnose, edit, verify, or restart. (arXiv) →read the paper

  • Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
    Shows that adding more environments can hurt unless diversity and difficulty are structured, then introduces ability-aware selection and hierarchical curricula. (arXiv) →read the paper

  • Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
    Forces agents to ground themselves across video frames before searching the web, reducing the tendency to bypass visual evidence in favor of easier text retrieval. (arXiv) →read the paper

  • Characterizing the Quality Profile of AI-Generated C++ in Production
    Tracks 3.52 million production code changes and identifies distinctive coupling, allocation, review, and compute costs in AI-generated C++. (arXiv) →read the paper

Post-training gets stricter about where supervision comes from

  • DAPD: Dual-Anchored Policy Distillation
    Identifies a privilege illusion in which students copy behavior dependent on teacher-only information, then aligns teacher and student along matched-information paths. (arXiv) →read the paper

  • When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
    Filters teacher signals that create large updates but remain weakly grounded in the input, preventing generic response templates from masquerading as useful evidence. (arXiv) →read the paper

  • ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
    Backtracks from a known answer to recover intermediate clues, then rewards useful search steps even inside failed trajectories. (arXiv) →read the paper

  • 🌟 SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
    Explains why multi-task supervised fine-tuning creates interfering updates while reinforcement learning tends to produce sparse, approximately orthogonal ones. (arXiv) →read the paper

  • Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
    Generates audio-grounded rubrics for each sample, then rewrites and reweights them as the policy improves so rewards continue targeting current weaknesses. (arXiv) →read the paper

Memory and context become reliability problems

  • 🌟 Addressable Memory for Video World Models
    Keeps compressed visual memory retrievable beyond the training horizon by assigning summaries in-distribution virtual positions rather than corrupting positional phases. (arXiv) →read the paper

  • 🌟 The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
    Shows that persistent-memory models routinely invent unsupported user attributes and that models reporting the least over-inference can be among the worst offenders. (arXiv) →read the paper

  • 🌟 SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
    Reveals that poisoned experiences can become durable skills, survive deletion of their source records, and evade checks that catch the original malicious trajectory. (arXiv) →read the paper

  • When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
    Shows that stale spatial memory can make an agent less safe than having no memory, particularly when visual auditing misses contradictions. (arXiv) →read the paper

  • Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
    Shows that relevance-based compression can retain an answer while deleting the references needed to interpret it, then restores missing dependencies with almost no additional context. (arXiv) →read the paper

World models separate prediction, verification, simulation, and control

  • Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
    Moves logged future trajectories from pre-decision hints to post-decision verification targets, reducing rationalization of outcomes shown to the model in advance. (arXiv) →read the paper

  • SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
    Uses video generation as a training-only dynamics signal, then discards the generator at inference so that a faster action model carries the learned world prior. (arXiv) →read the paper

  • MASS: Multiplayer World Models with Authoritative Shared State
    Separates a shared authoritative state from view-specific rendering, allowing many agents and cameras to inhabit one consistent learned simulation. (arXiv) →read the paper

  • WorldClaw: Agentic 3D Open-World Generation at Scale
    Decomposes open-world 3D generation into planning, terrain, assets, placement, and render-based repair, producing large scenes that remain editable. (arXiv) →read the paper

  • 🌟 Quo Vadis, World Modeling?
    Reframes world models as agent-centric interactive proxies that can return dynamics, execution results, memories, skills, or verification signals rather than only future physical states. (arXiv) →read the paper

Training and inference infrastructure becomes more economical

  • Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
    Maps how language, visual understanding, and generation help or compete during unified pretraining, finding early joint training more effective than late alignment. (arXiv) →read the paper

  • OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
    Compresses audio and visual tokens in two stages, preserving global structure before the language model and applying query-conditioned refinement after modalities interact. (arXiv) →read the paper

  • Skaling: Chinchilla’s Exponents Meet Kaplan’s Coupling
    Couples model capacity and data in one scaling law, improving extrapolation in data-scarce and heavily overtrained regimes with much smaller pilot sweeps. (arXiv) →read the paper

  • Modular TTT: Rethinking Test-Time Training as Composable Modules
    Decomposes test-time training into explicit, swappable modules so researchers can isolate which fast-weight architectures, losses, and update rules are actually helping. (arXiv) →read the paper

  • Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
    Moves teacher logits offline and fuses the KL loss so distillation runs faster, uses less memory, and supports longer contexts on a single GPU. (arXiv) →read the paper

Reply

Avatar

or to participate

Keep Reading

View more
caret-right