Today’s editorial: Has Google deliberately left the agent race? The evidence shows the opposite: Google is spending heavily to catch up, redirecting DeepMind’s researchers toward agents, and losing key scientists in the process. Google has not withdrawn. It is behind and internally divided.
This Week in Turing Post / Summer schedule:
Friday / AI Unicorns: Case for simulating 8 billion humans (Simile AI)
Sunday / Library: Frameworks Powering AI Agents
We recommend:
👨🔧 The Agent Harness – Why the LLM Is the Smallest Part of Your Agent System
Build the agent harness!
Prompts are the easy part. Production agents need state, tool coordination, observability, and evals to keep working.
Denouncing Bullshit: Why "The Actual Reason Why Google 'Fell Out' of the AI Race Changes Everything" Is Wrong
I've spent the last seven years analyzing the world of AI, and it makes me angry when fact-inverting speculation takes the air out of the room. I'm talking about The Algorithmic Bridge's article that even charitably I would call bullshit. Romero created 8,000 words, dense with links and evidence – and it's still bullshit, because he holds the evidence upside down. Why am I so agitated? Because it draws a completely false picture, and if you rely on it to think about Google, DeepMind, OpenAI, Anthropic, and the AGI race as a whole, you will be fooled. So let's concentrate on the same company Romero chose to "analyze," and see what's really going on there.
His hypothesis, “stated plainly”: Google didn't fall behind, it withdrew – Demis Hassabis doesn't believe in recursive self-improvement, he believes in world models, and he is "steering Google accordingly." Well, no matter how many of us would like that, Hassabis is not steering Google. Sundar Pichai, Sergey Brin, and Larry Page are. And every documented move they have made in the past six months points in the opposite direction from the one Romero describes. I’m also pretty sure Hassabis is highly interested in RSI, and I heard him saying that.
Romero states that Google "just “withdrew from agent race”. Ask yourself what a company that "withdrew" from the agent race would look like. Would it assemble an internal strike team – reported by The Information, with Brin's memo demanding the company "urgently bridge the gap in agentic execution and turn our models into primary developers" and framing the end goal as "AI takeoff, or AI that can improve itself"? Would its co-founder write an earlier memo declaring "the final race to AGI is afoot," pushing the Gemini team to become "the most efficient coders and AI scientists in the world by using our own AI," and prescribing 60-hour weeks? Would Page, per reporting that has circulated since 2024, tell people inside Google he is willing to go bankrupt rather than lose this race? Would its CEO announce, on the Q2 earnings call, the company's largest training run ever for Gemini 4, defend capex guidance of $195–205 billion, and let contractual commitments climb past $800 billion – enough to push Alphabet into its first negative free cash flow quarter since the IPO? Romero cites several of these facts himself. He quotes the strike team and Page's bankruptcy line. But he concludes withdrawal anyway, because – I don’t know – maybe the facts are inconvenient for the headline and the headline was written first? That’s far from good analysis.
What Romero does get right is that Hassabis and Pichai want different things from AGI. And then he inflates one man's research convictions into corporate strategy – and misreads the convictions themselves. Read the actual interview. Yes, Hassabis says world models are "the thing I'm mostly spending my research time on" and answers "I do" when asked about a ChatGPT moment for them. But in the same breath he lists his other main research areas: "extensions in science that we're doing with like AlphaEvolve," and – fatal to Romero's framing – "agentic stuff and improving memory, reliability." He calls world models "critical components in getting to AGI" that "make use of Gemini under the hood," and calls Gemini itself "a key component of what we see as an eventual AGI system." Components, plural, stacked on the LLM – not an exit from it. He even describes his own time as "about half and half" between research and "turbocharging the app," with DeepMind as "the engine room of Google." That is not a man steering the company toward a contrarian bet; that is a scientist with a broader theory of intelligence describing, in his own words, how much of his life now goes to the product war. His DeepMind was built to solve grand scientific challenges. Pichai and the founders want to beat Anthropic at enterprise agents, now. Both visions live inside one company, and only one of them controls the budget.
You can watch who is winning that argument in the personnel record. Google moved John Jumper – its Nobel laureate, the living symbol of the science-first DeepMind – onto the Code Strike coding team, conscripting him into the coding-agent catch-up war. Within months, Jumper left for Anthropic, taking Jonas Adler and Alexander Pritzel with him. The FT reports the AlphaFold team has been disbanded outright: most original authors reassigned to Gemini projects, nearly a quarter gone from the company entirely – and DeepMind's own research VP Pushmeet Kohli confirmed the shift on the record: "The strategy has evolved." Noam Shazeer, a $2 billion acquihire and Gemini co-lead, walked to OpenAI. Alphabet's steepest single-day decline in over a year followed, with reported estimates of roughly $270 billion in market value erased across two sessions. Reporting on the exits keeps returning to compute conflicts – the exact signature of two visions fighting over one pool of resources. Meanwhile Gemini 3.5 Pro slipped months behind schedule after a training-data refresh produced disappointing results, a securities-fraud investigation opened, and an employee told Axios the thing no "strategic withdrawal" narrative survives: "We're behind."
They are behind. That is the truth – not that they have withdrawn. A comment from Logan Kilpatrick under Mark Cuban's share of Romero's article is interesting here:
So no, Google did not "fall out" of the AI race by choice, and no, Hassabis has not converted Alphabet to the world-model gospel. What we are watching is far more consequential: the corporate leadership dragging the company into a head-on fight on OpenAI's and Anthropic's terms, the chief scientist's life's work being stripped for parts to feed it, and the best researchers voting with their feet on the way out.
What we have is this: Anthropic has doctrine. OpenAI has a center of gravity. Google has a disagreement between the man who runs its research and the men who run its money – and it is spending $200 billion a year funding both sides of it. This is what facts tell us.
Would Google be better off if Demis Hassabis ran its AI efforts end to end? Perhaps, but that is not happening. Would Hassabis be better off leaving to pursue AI science free from corporate constraints? Considering the recent wave of investment, this might be the best possible time for him to do so.
I, personally, would love that.
If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going.
📹 The End of the Mathematical Guild? In this episode of Attention Span, we discuss what math looks like now when proofs are no longer a bottleneck. Very interesting times! Check it out →
News from the usual suspects ™
OpenAI: academics get the frontier
OpenAI launched ChatGPT for Academic Researchers, offering free access to GPT-5.6, Codex, ChatGPT Work, larger context windows, and more than 75 life-science tools.
The program starts with 10,000 researchers this summer and is expected to reach 100,000 scientists, mathematicians, and engineers through 2027. OpenAI has committed more than $250 million to external scientific research.
Europe: gigafactories on its mind
The European Union announced €10 billion in public funding for seven AI gigafactories, up from the five originally planned.
The Commission wants to attract another €20 billion from private investors. The facilities will combine processors, software, cloud infrastructure, networking, and data centers, supplementing 19 existing European AI factories.
AMD, NVIDIA, and Qualcomm have signed letters indicating their willingness to supply chips. Europe wants technological sovereignty and has concluded that sovereignty requires a procurement process.
Google: Earth lasts one day
Google introduced AI image generation inside Google Earth, allowing users to create synthetic scenes grounded in real satellite imagery and three-dimensional locations.
It paused the feature one day later after users produced content that appeared to violate its policies.
The images were watermarked and separated from Google Earth’s main experience. But synthetic events become considerably more persuasive when placed inside recognizable real locations.
MiniMax: video becomes fully multimodal
MiniMax released H3, an open multimodal model that accepts text, images, video, and audio as shared context, then generates videos with native stereo sound.
H3 supports clips of up to 15 seconds at 2K resolution and can use one video as a motion reference, an image as the subject, and a separate audio file for sound or vocals.
MiniMax says 2K generation costs less than one-third as much per second as mainstream competitors and plans to release the weights, subject to regulatory restrictions.
Thinking Machines: Inkling gets smaller
Thinking Machines released Inkling-Small, an open-weight multimodal model with 276 billion total parameters and 12 billion active parameters.
It accepts text, images, and audio, supports a one-million-token context window, and is available under the Apache 2.0 license. A quantized version can run with 180GB of GPU memory, compared with at least 600GB for the full-precision checkpoint.
DeepSeek and Alibaba: cheap and large
DeepSeek officially released V4-Flash, with stronger coding, tool-use, and agent capabilities and native support for the Responses API used by Codex-style systems.
It charges $0.14 per million input tokens and $0.28 per million output tokens. Artificial Analysis estimated an average cost of three cents per benchmark task, compared with $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5.
Its performance is lower than the leading OpenAI and Anthropic models, but roughly comparable to Google Gemini 3.6 Flash. DeepSeek’s argument remains disciplined: capable enough, much cheaper.
Alibaba answered from the other direction with Qwen3.8-Max, its largest model yet, at 2.4 trillion parameters.
The mixture-of-experts model activates 95 billion parameters per request, supports text, images, video, and one million tokens of context, and became the highest-ranked Chinese text model on Arena.AI. It ranked second globally for visual understanding, behind a Claude Fable 5 variant.
China produced both ends of the market in one weekend: one model optimized for scale, another for price.
Other interesting models
Alibaba: one agent for every interface
Alibaba introduced Qwen-UI-Agent, designed to operate across phones, browsers, desktops, terminals, search engines, and APIs.
It can combine clicks, commands, and tool calls while preserving task state across environments. Alibaba trained it using thousands of sandboxes and more than 100 physical phones.
Microsoft: video models process only what changes
Microsoft released Mage-VL, a streaming multimodal model that uses motion vectors and residuals from compressed video instead of repeatedly processing complete frames.
Microsoft reports more than 75 percent fewer visual tokens and up to 3.5 times faster inference.
Mistral AI: moderation adapts to the policy
Mistral introduced Shieldstral, a three-billion-parameter multimodal safety model that evaluates content against natural-language rules.
One model can therefore support different products, jurisdictions, and definitions of harmful content without retraining a separate classifier.
MemTensor: memory moves inside the model
MemTensor released Metis in 4-billion, 9-billion, and 27-billion-parameter versions.
Metis maintains a persistent internal memory state that updates during inference without changing the model’s weights or replaying the full conversation.
Frontis AI: a model improves ML systems
Frontis AI released Frontis-MA1, a 35-billion-parameter model trained to draft, debug, improve, and combine machine-learning solutions.
Its proposals can be executed and evaluated inside a controlled environment, creating a measurable loop for iterative improvement.
CASIA: world models learn a physical language
PhiZero represents physical changes as discrete tokens before rendering them back into video.
The design separates dynamics from visual appearance, making predicted actions and transitions easier to transfer across objects and environments.
Huawei and HUST: robot control loses the LLM
TurboVLA is a 200-million-parameter vision-language-action model that predicts robot movements without routing every decision through a large language model.
It runs at 32 updates per second on an RTX 4090 while using less than 1GB of memory.
Research
Trends we see looking at every paper related to AI and ML published last week:
Agents begin improving the tools, skills, and rubrics around them
Evaluation becomes part of the training loop
Memory separates from both model weights and context windows
Retrieval becomes an execution policy, not just a ranking step
Evidence and provenance become persistent agent infrastructure
World models move from predicting actions to controlling them
Video generation gets faster without surrendering diversity
Open-ended research exposes scientific judgment as the remaining bottleneck
World models, embodied action, and efficient generation
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Shows that sufficiently accurate robot-free demonstrations may eliminate the real-robot post-training normally required before deployment. This could shift manipulation-data collection away from expensive robot fleets toward scalable human demonstration systems. →read the paper🌟 Wonder: Video World Model Done Better
Constructs interactive video worlds that users can navigate, revisit, and explore over long horizons. Its combination of camera control, persistent memory, and real-time generation addresses the difference between producing a convincing clip and maintaining a coherent explorable world. →read the paperShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
Learns transferable actions from paired videos that preserve dynamics while changing visual appearance. This allows a demonstrated behavior to become a reusable control signal across new subjects and environments without requiring structured action labels. →read the paperINTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
Learns a direct mapping from a desired visual change to the action needed to produce it. The approach removes the expensive test-time search usually required to turn a predictive world model into a practical controller. →read the paper🌟 DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Explains why accelerated video generators can preserve visible quality while quietly losing diversity. It aligns mode-covering and mode-seeking objectives so that distillation retains more of the teacher model’s underlying distribution rather than collapsing toward a narrow set of attractive outputs. →read the paper🌟 Parallel Decoding Distillation for Fast Image and Video Generation
Accelerates diffusion and flow models by predicting several denoising transitions within each model evaluation. It offers a simpler alternative to adversarial and variational distillation methods, whose training instability frequently produces motion loss and mode collapse. →read the paper
Agent learning, self-improvement, and reliability
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Co-evolves the instructions used to solve a task and the rubric used to evaluate the solution, while preventing improvement from being faked by quietly making the rubric easier. It provides a more credible mechanism for inspectable self-improvement in open-ended tasks. →read the paperCoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Converts rubric feedback into token-level credit by comparing the same response with and without the rubric context. This lets reinforcement learning identify which parts of an answer actually satisfied the criteria without training another token-scoring model. →read the paperSkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Trains a single policy to solve related tasks while continuously updating a reusable skill document. The important shift is from learning separately inside each episode to accumulating procedures that improve performance on later, different tasks. →read the paper🌟 GPT-Red: Automated Red Teaming via Self-Play at Scale
Scales prompt-injection discovery through attacker-defender self-play and uses the resulting attacks to adversarially train stronger models. It turns red teaming from a periodic external audit into a continuously escalating training loop. →read the paper🌟 Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies
Tests autonomous agents on genuinely open-ended research rather than narrowly verifiable tasks. It finds that agents can complete much of the engineering work but still struggle with scientific judgment, backtracking, creativity, and deciding what constitutes a meaningful contribution. →read the paper
Memory, retrieval, and evidence infrastructure
🌟 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Changes the unit of scientific retrieval from entire papers to atomic claims carrying quotations, evidence locations, and provenance. This gives research agents a more reliable substrate for synthesizing findings across publications without losing the connection to source evidence. →read the paperCodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
Treats repository context as maintained infrastructure rather than something every coding agent must rediscover from scratch. It combines lexical, semantic, and structural views that persist across code changes and can be served within bounded context budgets. →read the paper🌟A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
Uses relevance scores to determine where an agent begins searching, which documents it explores, and which local evidence remains visible. It turns relevance from a one-time ranking signal into an execution policy for navigating large corpora. →read the paperMemory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Separates memory capacity from the reasoning backbone and scales a dedicated parametric memory module independently. The results suggest that adding parameters to memory can be more efficient than making the complete language model larger. →read the paperLEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Constrains an agent’s reasoning to evidence explicitly recorded by its tools. Every downstream claim must cite an active ledger entry, while repair operations are prevented from introducing unsupported content, making provenance part of the agent’s execution state rather than an afterthought. →read the paper
How did you like it?
/









