This website uses cookies

Read our Privacy policy and Terms of use for more information.

Today’s editorial: The superintelligence we are racing toward – and how to record our path to it.

We recommend:

🛠 Move from AI Investment to Outcomes

Building AI applications at scale requires solving fluid challenges across models, agents, and infrastructure. AWS offers capabilities at every layer of the stack with a partner ecosystem enabling ISVs to build and operate AI more efficiently at scale.

Access hundreds of foundation models through Amazon Bedrock, and explore over 4,000 AI tools in the AWS Marketplace.

Permanent Dawn

I want to stop. I want to take a few steps back actually, because I need to see the whole picture and I have not been able to.

We are racing – through the spasms and the fever of social media, through the launches and the leaderboards and the week's argument about timelines – toward something that none of us has managed to describe. What surprises me most is that even science fiction, which used to be reliable for a glimpse of what was coming, is no help here.

What is this superintelligence we are at the dawn of? The disagreement about when it arrives seems to me the smaller trouble. What bothers me is that we are reorganizing our lives around a thing we have not yet put into words.

There is an old way of finding out what a person is made of, and it is always the same method: take the help away and see what remains. It is how we test a student, how a craft decides an apprentice is finished, how a parent notices a child has grown. You withdraw the instructions, and whatever is still standing afterward is the person.

A benchmark ASI-Bench: At the Dawn of Artificial Superintelligence published last week – the thing that started me thinking about the dawn of superintelligence in the first place – runs that experiment on machines. Sixty research projects across eleven sciences, served at four levels of help: full procedure, then only the name of the method, then only the goal and the data. The scores fall off a cliff, and they fall at a particular place, the moment the written procedure is removed. Take away the steps and most of the capability goes with them. Take away everything else and little more is lost, because there was not much else there to lose.

That is the pretext, and it is only a pretext. The question underneath it is a great deal older than the technology.

The part that cannot be written down

Michael Polanyi, the Hungarian-British polymath, gave this a name in 1966 (in The Tacit Dimension): we know more than we can tell. The surgeon's hands. The editor's ear for a sentence that has gone false. The scientist's suspicion that a result is too clean. Little of it survives transcription. It passes by standing next to someone for years, and it tends to die with people who never had an apprentice.

I want to be that apprentice. Standing next to a thing for long enough to catch what it cannot say about itself strikes me as a reasonable job description, for a person and for a publication both.

We rehearsed for a different arrival

We have not managed to describe this thing, and yet we spent a century describing it in advance.

We have "seen" the robot with a body, countable and discrete, standing in a doorway. We were persuaded it would be a hostile singular mind with a plan of its own. There was an android asking to be recognized as a person, and others besides – most of those stories assumed the machine would want something.

What arrived has no body and no edges, and wants nothing of its own. It is not singular and not continuous, and it does not persist between conversations. It is not hostile, and its characteristic failure is not rebellion but a fluent, untroubled wrongness that few novelists thought to invent. It came through a text box, priced like a streaming service. Or even for free.

It also took the wrong things first. The tradition assumed arithmetic and heavy lifting would fall early, and that poetry, argument, and drawing would be the last human ground. The so-called Moravec paradox, that we debunked a couple of years ago.

There were people who saw some angles of what we have. E.M. Forster wrote "The Machine Stops" in 1909, about people who live alone in cells and consult a disembodied system through a screen for everything, including company (but we do not live in cells). Stanisław Lem's Golem XIV is a superintelligence that lectures its audience and has no particular interest in them (but it has interests of its own). In 2013, Her gave us a voice-first, bodiless, emotionally competent system sold as a consumer product (but Her still felt too human-like, and LLMs are not).

Image Credit: Warner Bros

None of them works as the image for what we have now, or for what is coming. Fiction was never forecasting. It was rehearsal – the advance picture that tells people where to stand when the thing walks in. We rehearsed for the uprising and the rights hearing. We did not rehearse for a colleague with no self, who writes better than we do and is sometimes confidently wrong about exactly the things we are least able to check. And we certainly have not rehearsed for abundance that somehow became associated with that very text box.

We have too many images of machine intelligence, most of them wrong, and now that wrongness is fogging the real picture.

Who has an audience

The people who asked the larger question well were not the ones imagining machines, and I think that is why they lasted.

Keynes asked it in 1930. He guessed the economic problem would be solved within a century, and then, instead of celebrating, he worried. He thought we would be delivered into our permanent problem – how to occupy a life that necessity no longer occupies – and he expected something like a collective nervous breakdown, judging by the wealthy women of his own time who had been released from need and found little on the other side.

Arendt asked it in 1958. The opening pages of The Human Condition describe a society of laborers about to be freed from labor, which she thought close to the worst thing that could happen, because such a society knows of nothing better and has nothing else it knows how to do. Her separation of labor from work from action remains, for me, the most useful equipment anyone has built for this moment.

Bernard Suits asked it in 1978 and gave the strangest answer. If every instrumental activity became unnecessary, what would be left is games – voluntary attempts to overcome unnecessary obstacles – not as a consolation prize, but as the highest form of existence available to a being with nothing it has to do.

They lasted because each described the human condition without necessity, and that description does not depend on what the machine turns out to be. The novelists specified the hardware, and the hardware is what rotted. Abstraction outlived imagination, which is not the usual result.

People are working on this now – Shannon Vallor on what these systems reflect back at us, John Danaher on automation and utopia, Elizabeth Anderson on work and freedom, Michael Sandel on merit and dignity, Kieran Setiya and Susan Wolf on meaning. The problem is that the asking has no audience where the decisions are made. It is not flashy, it does not sell an LLM or a world model, and it will not be reposted by Elon Musk. And if you think about it, the most important topics are currently discussed on X, which is essentially a living feed. It is a remarkable way to stay inside the Silicon Valley bubble and read every mover in the industry at once, but even their words turn elusive there, because they dissolve into the noise of everyone else's.

What I want to do, and what I want to ask you about

Here is the thing I have been circling for months, and I would like your advice before I commit to it. (That was the topic I wanted to discuss with you last Friday, but I couldn’t formulate it yet.)

I do not think we need to predict the future, which is what so many reports spend their pages doing. I think we need to register it carefully – write down what was claimed, by whom, and when – and then go back and check at each stage. Kept up for long enough, that unfolds the future for us without anyone having to forecast anything.

So I want Turing Post to keep a register. Not forecasts, and not another feed of takes, but a running record. What was claimed. Who claimed it. What would have to be true for the claim to hold. And then, at intervals, what happened. The claims themselves are easy to find and easy to forget, which is the whole trouble, because they are made in a format designed to be forgotten. A register – a ledger, am almanac? – would hold them still long enough to be checked.

Part of that belongs online, where it can be corrected and extended. But I have come to want a material version as well: something printed, arriving a few times a year, something extremely beautiful that you can put on a shelf and take down in 2030 to see what we believed in 2026 and how much of it survived. Or even read it to your children. They often see things we miss.

And yes, I just need to hold it in my hands and flip the pages, don’t you?

Alongside the record I want the reflection, which is where the philosophers and the economists come in. Not commentary on the week, but people willing to say what a claim would mean for a life rather than for a valuation. I have been building toward this with the economists already. The philosophers are the next step.

What I do not know is where the line sits, and this is the part I would like help with. Whether a register is something you want at all, or whether the weekly explanation is enough. Whether print reads as serious or as nostalgia. What are we missing in this whirlpool of news and changes?

I am not confident in any of this. I am more confident that we are describing the wrong problem. We are doing it very enthusiastically and very loudly, but I keep thinking we are missing the bigger picture.

So I would like to know what you see. Send me your thoughts. I do not trust a poll on that.

You can reply to this email or –

P.S. Turing Post has always tried to connect the development of AI with the humans building it and living with it. But some questions are too large for the daily news cycle. They need time, history, disagreement and repeated examination.

Each quarterly Almanac could take one such question and follow it across technology, economics, institutions, history and human life. Not to manufacture a final answer, but to understand the choices being made while those choices are still ours.

If this really is the dawn of superintelligence, we should document more than how intelligent the machines become.

We should ask what kind of humans we intend to be beside them.

📹 Let’s Discuss Ox Alpha – the Model That Turned Us Into the Benchmark. Check it out this video →

Follow us on

We are reading / watching

News from the usual suspects ™

Deployment, routing, payments, and infrastructure. Such was the week.

  • OpenAI: teens, privacy, and lower prices

    OpenAI launched ChatGPT for Teens, automatically placing users aged or estimated to be 13–17 into an experience with stronger safeguards, parental controls, Study Mode, homework reminders, and scheduled study hours.

  • SpaceXAI: distribution widens

    Grok 4.6 arrived on Amazon Bedrock and Google Enterprise Agent Platform. Grok Build expanded to every plan on web and mobile, while Grok Bot became available through more Cursor subscriptions.

    This was a distribution week for SpaceXAI: the model entered both major clouds while its agents spread across the company’s own developer products.

  • Mistral: search and sovereignty

    Mistral launched Agentic Search, a retrieval layer that lets AI systems search, open, navigate, read, and verify information across complex document collections. It is available through Mistral’s Search Toolkit and Libraries.

    France also said it would use sovereign AI providers such as Mistral, rather than OpenAI, to test government services for cybersecurity vulnerabilities after a tax-agency breach. Sovereign AI is becoming a procurement decision.

  • Alibaba: cloud growth requires fresh capital

    Alibaba’s AI Cloud and Compute revenue rose 45%, while quarterly profit fell 75% and capital expenditure increased 75%. The company has already spent nearly half of its planned $56.4 billion AI investment through 2029.

    Alibaba then launched a $10.2 billion share placement, directing the proceeds toward chips, infrastructure, and models, and released its Wan3.0 video-generation model. The AI business is growing quickly; so is the financing requirement.

  • NVIDIA: AgentX measures the real workload

    NVIDIA published the first on-silicon Vera Rubin NVL72 results on AgentX, SemiAnalysis’s benchmark built from recorded, multi-step coding-agent sessions.

    In NVIDIA’s measurements, Vera Rubin delivered up to 30 times more throughput per megawatt and 35 times lower token cost than GB300 NVL72 on DeepSeek V4 Pro. The Vera results are still pending SemiAnalysis review. Existing AgentX data shows GB300 delivering up to 15 times more throughput per megawatt than H200.

  • NVIDIA at Hot Chips: the platform broadens

    NVIDIA put Groq 3 LPX into full production. Artificial Analysis measured 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, with Nebius becoming the first cloud provider to adopt it.

    SpaceXAI will use NVIDIA Vera CPUs for agent orchestration, tool use, code execution, and data processing, and plans to base its first Starmind AI satellite on an optimized Vera Rubin NVL72 system.

    CoreWeave has deployed Spectrum-X Multiplane in production. NVIDIA also introduced Scale-In infrastructure for networking, storage, security, and observability, while NVLink Fusion connects custom CPUs and XPUs to NVIDIA’s rack, networking, cooling, and management stack. The company is extending its position from accelerators into the full operating architecture of large AI systems.

  • Etched raised $700 million at a $21 billion valuation, more than doubling its valuation in less than a month.

    Jane Street led the round, became Etched’s first customer, and began deploying the company’s first rack. The valuation is aggressive, but the hardware has reached a customer’s data center. We made a video about it:

Models

  • Meta presented MoE-ViE, a family of mixture-of-experts vision encoders that routes visual tokens through specialized experts.

    Its video training uses frame-level distillation and selective expert freezing to add efficient temporal understanding without losing the image capabilities learned during pretraining.

  • AgiBot Finch and Shanghai Innovation Institute presented τ0-VLA, a hierarchical robot foundation model that spends additional test-time compute on difficult high-level decisions.

    A world model predicts the likely outcomes of possible plans before a lower-level vision-language-action policy executes them.

  • The teams presented 4DAnyone, which reconstructs an animatable human in four dimensions from one casual, uncalibrated video.

    It generates consistent alternative viewpoints and lifts them into a dynamic Gaussian representation that can be viewed from new camera angles.

  • Tencent Hunyuan presented WithEveryone, an image model designed to preserve multiple supplied identities inside the same generated scene.

    It explicitly plans where each person should appear before rendering, then grounds identity features to those regions to reduce blending and accidental duplication.

  • OPPO Research Institute and Hong Kong Polytechnic University presented PixRestore, a compact diffusion transformer for multiple image-restoration tasks.

    It works directly in pixel space without a VAE or text-to-image pretraining and can be distilled into a one-step generator.

Research

Trends we see looking at every paper related to AI and ML published last week:

  • Agents learn through harnesses, environments, and reusable skills

  • Verification and governance move inside the workflow

  • Reasoning, memory, and post-training become adaptive

  • World models connect directly to physical action

  • Scaling, serving, and generation become more economical

  • AI systems move upstream into discovery

Agents learn through harnesses, environments, and reusable skills

  • Agent Lightning v1.0: Towards Harnessed Agentic RL
    Connects arbitrary deploy-time agent harnesses to reinforcement learning without requiring the trainer to own or rewrite their execution loops. →read the paper

  • LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
    Brings policy-gradient training into native coding harnesses while addressing sandbox failures, reward hacking, and differences between training rollouts and real agent execution. →read the paper

  • ClawGym II: Exploring Black-Box RL on Agent Harness
    Reconstructs opaque multi-turn agent trajectories as prefix trees, allowing reinforcement learning to work across heterogeneous harnesses that expose only inputs and outputs. →read the paper

  • Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
    Replaces backpropagation-heavy agent reinforcement learning with evolution strategies that learn from complete trajectory rewards using inference-level GPU memory. →read the paper

  • EnvHarness: Awakening Static Worlds for Agent Learning
    Wraps static environments in programmable components that generate new tasks around diagnosed policy weaknesses while preserving the original environment and its verifiers. →read the paper

  • SPADE: Self-Play in Adaptive Synthetic Executable Environments
    Co-trains one model as both environment designer and reasoning agent, allowing executable training worlds to become harder and more useful as the policy improves. →read the paper

  • AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
    Synthesizes persistent business worlds containing entities, tools, services, state, and executable rules from which many different agent tasks can emerge. →read the paper

  • SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
    Synthesizes repository-specific problems and distills successful repairs into reusable skills before the agent encounters real issues in the project. →read the paper

  • SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
    Separates credit for choosing a skill from credit for executing it, preventing a correct selection from being punished because the later trajectory failed. →read the paper

  • FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
    Builds terminal tasks around a shared executable state so that instructions, reference solutions, environments, and verifiers remain mutually consistent. →read the paper

  • Looped Language Models Improve Compositional Tool Calling
    Uses recurrent computation to improve multi-step tool coordination, intermediate-state tracking, and preservation of dependencies between calls. →read the paper

  • Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
    Frames multi-agent orchestration as the construction of evolving graphs over tasks, specialized agents, dependencies, verification steps, and persistent execution state. →read the paper

Verification and governance move inside the workflow

  • SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
    Declares industrial control code complete only after specification checks, compilation, and live-runtime verification have all succeeded. →read the paper

  • PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
    Compiles organizational policies into workflow graphs that actively guide agents through required procedures instead of merely blocking individual forbidden actions. →read the paper

  • CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
    Routes a safety adapter continuously from the model’s hidden states, applying stronger intervention to risky prompts while preserving ordinary behavior on benign requests. →read the paper

Reasoning, memory, and post-training become adaptive

  • Chain-of-Experience for Continual LLM Improvement
    Accumulates feedback and experiential traces across repeated test-time attempts, turning inference into a continual improvement process rather than a sequence of independent queries. →read the paper

  • Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
    Derives reward signals from a diverse cohort of peer models, reducing dependence on ground-truth labels and the correlated errors produced by self-reward. →read the paper

  • Learn What’s Left, Not What’s Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
    Redirects gradient budget away from objectives the policy has already mastered and toward rewards with meaningful remaining headroom. →read the paper

  • Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
    Shows that on-policy distillation transfers reasoning behavior broadly when teacher and student share a model lineage, but generalizes less reliably across unrelated families. →read the paper

  • ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
    Allocates parallel reasoning compute according to whether each branch’s tentative answers are converging, allowing weak or redundant branches to stop early. →read the paper

  • Cross-Model Memory Transfer via Target-Side Reader Adaptation
    Transfers frozen external memories between different model families by adapting only a lightweight reader attached to the target model. →read the paper

  • Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
    Separates document injection, question-answer alignment, and general-ability recovery so that a bounded corpus becomes usable parametric knowledge without retrieval at inference. →read the paper

World models connect directly to physical action

  • Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
    Evolves runtime critics and recovery skills around a frozen robot policy while physical execution is still unfolding. →read the paper

  • Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
    Aligns latent geometry with actual decision progress, showing that a representation can reconstruct the world accurately while still producing a poor planning cost. →read the paper

  • EXIMO: VLM Guided Exploration of VLA Policies
    Combines vision-language-guided exploration, imitation learning, and residual reinforcement learning to adapt large robot policies with less human teleoperation. →read the paper

Scaling, serving, and generation become more economical

  • FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
    Treats CPU, GPU, memory bandwidth, and expert residency on a personal computer as one adaptive platform for serving large mixture-of-experts models. →read the paper

  • FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
    Moves block-sparse prefill attention toward production through corrected approximation methods, specialized kernels, and integration with serving systems. →read the paper

  • Let’s Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
    Predicts full-scale mixture-of-experts learning rates from small proxy models by transferring across model width and extrapolating across token budgets. →read the paper

  • Abra: Scaling Diffusion Image Training
    Establishes predictable scaling relationships for text-to-image diffusion and finds that compute-optimal visual generators require substantially more training data per parameter than language models. →read the paper

  • From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
    Organizes image-generation data and curriculum around dependencies among grounding, editing, and knowledge capabilities instead of treating each training corpus independently. →read the paper

  • InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
    Extends instruction-based editing to continuous video streams whose future frames and edit requests arrive incrementally rather than being available in advance. →read the paper

  • Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
    Accelerates video and world-model transformers by sparsifying attention and reconstructing omitted responses from a small number of exact probes, without retraining the underlying model. →read the paper

  • Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
    Compresses an eleven-billion-parameter vision-language model into a format small enough for mobile hardware using self-generated calibration data and CPU-oriented quantization. →read the paper

AI systems move upstream into discovery

  • Improving the Matrix Multiplication Exponent with Modern Optimization and AlphaEvolve
    Reformulates the optimization behind the best matrix-multiplication bounds and combines modern solvers with evolutionary search to improve the known upper bound. →read the paper

  • The Problem Is the Problem: Towards Scalable Mathematical Discovery
    Moves AI-assisted mathematics upstream by searching the literature for open problems, attempting them, and triaging the most promising results for expert review. →read the paper

FAQ

What is Artificial Superintelligence (ASI)?

Artificial Superintelligence describes hypothetical AI systems that surpass human cognitive capabilities across virtually every economically valuable domain, including scientific research, strategic planning, and complex problem-solving.

What is the difference between AGI and ASI?

Artificial General Intelligence (AGI) denotes systems matching human-level competence across broad tasks, whereas Artificial Superintelligence (ASI) describes systems operating far beyond human capabilities, particularly in unassisted scientific reasoning, autonomous discovery, and complex orchestration.

What is the tacit dimension in AI capability?

Coined by philosopher Michael Polanyi ("we know more than we can tell"), the tacit dimension refers to intuitive human knowledge that resists explicit codification. Current frontier models excel at step-by-step instructions but struggle when written procedural guidance is removed, exposing this missing procedural intuition.

How do scientific benchmarks measure the path to superintelligence?

Modern benchmarks test model autonomy across multiple sciences by progressively withholding explicit instructions. These evaluations reveal that while models can execute known procedural paths with high accuracy, independent hypothesis generation and unguided scientific exploration remain major hurdles.

Why is an AI claim register or almanac or ledge necessary?

Tracking explicit technical predictions, their underlying assumptions, and retrospective real-world outcomes creates a verifiable historical record of progress, distinguishing tangible milestones from shifting industry marketing.

Reply

Avatar

or to participate

Keep Reading

View more
caret-right