Join me for a closer look at agent memory
We’ve partnered with O’Reilly to explore a question that keeps coming up as AI agents take on longer, more complicated work: what should they remember, and how do we know that memory is helping?
On October 13, I’ll be chairing the free, virtual Expert Showcase:
Agent Memory
We’ve curated the lineup to bring together people approaching this problem from very different directions, and I’d love to have you join us.
You will listen to Pete Johnson, Lewis Liu, Shawn Shen, Jason Davenport, Jess Weng, Kris Murphy, Xavier Bull, and Sally O’Malley, who are bringing perspectives from MongoDB, Microsoft, Memories.ai, Google Cloud, LangChain, NVIDIA, Adarga, and Red Hat/OpenClaw.
We’ll look at how agents remember without keeping everything in context, how reflection can turn past experience into useful memories, and how to measure whether an agent actually improves. We’ll also explore visual memory for physical agents, memory in high-stakes settings, and the trust boundaries that determine what an agent should retain or forget.
Bring the questions you’re working through. We start at 11 a.m. ET / 8 a.m. PT.
Share Turing Post with one person. You will help us grow.
And now to the new episode in our AI Builds AI series:
Skill Up: How AI Builds the Skills Other AI Uses
What should a coding agent leave behind after it finishes a job? Simple question, right? Working code, hopefully, and evidence that it works. But the session may also reveal a few details that can be useful for the next task.
And we can actually ask an agent to examine that experience and suggest improvements to its working environment. That’s exactly what Matt Pocock’s /retro skill does, and that’s inspired me to explore this topic. Matt has one particularly sensible instruction: to turn mechanically detectable mistakes into automated checks. Reserve written guidance for decisions that require judgment.
That sounds like a good start to go deeper into the new wave of enthusiasm for skills. Turning an agent’s experience into a procedure, test, tool, or environment change takes engineering judgment. And researchers are now trying to automate that choice. There is plenty of them: Microsoft has SkillOpt. Renmin University and Tencent have SkillAdam. OpenHands has SkillRefiner. Alibaba has skill-up. Anthropic and OpenAI have their own takes on skills.
What interests me is how agents turn completed work into procedures other agents can use: thats’s AI Builds AI in practice.
But what makes a lesson reusable? And how do we prevent a perfectly sensible correction from becoming tomorrow’s mistake?
In this episode, we discuss:
What an agent skill actually contains
Why readable instructions can still be difficult to use
How SkillOpt and SkillAdam improve a procedure
What SkillRefiner can learn without replaying the work
How to tell whether a skill generalizes
When a skill should move into code or model weights
A development workflow you can apply to your own agents
With examples from Microsoft’s SkillOpt, Renmin University and Tencent’s SkillAdam, Sakana’s ShinkaEvolve, OpenHands’ SkillRefiner, and Alibaba’s skill-up, and many more.
What are we putting in a skill
Some basics first: a skill can be as simple as instructions in a SKILL.md file, accompanied by scripts, examples, and references. Its description helps an agent decide when to load it. The procedure then supplies knowledge about how to carry out a particular kind of work.
Here is a miniml skill.md example:
---
name: explain-code description: Use when the user asks to explain a function or code snippet.
---
Explain what the code does in plain language.
1. State its purpose in one sentence.
2. Describe its inputs and outputs.
3. Walk through a small example.
For an agent reviewing changes to an ML pipeline, that might include where the evaluation data lives, which checks detect leakage, how to compare results, and what evidence belongs in the review. The model may already understand Python and experimental design. When it’s an agent at work, the model needs to discover how those things are implemented here.
Anthropic makes a distinction between skills that address gaps in a model’s capabilities and skills that encode a team’s preferred process. Its March update to skill-creator added evaluation and benchmarking support, including pass rates, elapsed time, and token usage. That lets authors investigate whether a skill helps, whether an edit broke it, and whether a newer model still needs it.
So a skill is not a particularly elaborate prompt. A skill is a reusable procedure that makes a claim about what will remain true across future tasks. The principles and required formats generalize across tasks; the paths and specific commands belong strictly to the repository.
For readers of our MCP guide, the relationship is straightforward: an MCP connection can expose experiment records or repository tools. The skill can explain how to use those capabilities for a task. Access, procedure, and evidence of success are separate things to design.
Once you write the procedure down, however, you encounter a surprisingly basic problem: how will the agent know that this is the right one?
Don’t settle for shallow articles. Learn the basics and go deeper with us. Find inspiration for what to build and the knowledge to put it into practice.
Join Premium members from top companies like Microsoft, Nvidia, Google, Hugging Face, OpenAI, a16z, plus AI labs such as Ai2, MIT, Berkeley, .gov, and thousands of others to really understand what’s going on in AI.
How did you like it?
Here is a simple table when you think about all the papers and repositories we shared:
Approach | What happens | Examples |
|---|---|---|
Learn from new runs | Run tasks, examine failures, edit the skill, and test again. | SkillOpt, SkillAdam |
Learn from past work | Analyze recorded runs and outcomes to propose better instructions. | SkillRefiner, CASD |
Learn from existing code | Turn repository knowledge into procedures other agents can use. | Repo-To-Skill |
Check whether skills help | Compare performance with and without skills, including on unfamiliar tasks. | Alibaba’s skill-up; held-out generalization research |
Train the procedure into the model | Use skill guidance during training, then gradually remove it. | SKILL0 |
Frequently asked questions
What is an AI agent skill?
An AI agent skill is a reusable procedure that tells an agent how to carry out a particular kind of task. It can consist of a SKILL.md file with an activation description and instructions, plus supporting scripts, examples, and references.
How do AI agents improve their own skills?
Agents can examine execution records and outcomes, identify recurring problems, and propose changes to their instructions. SkillOpt and SkillAdam test revisions through new runs, while SkillRefiner proposes edits from historical traces. A proposed improvement still needs evaluation on separate tasks.
What is the difference between an AI agent skill and an MCP connection?
An MCP connection gives an agent access to tools or data. A skill explains when and how to use those capabilities to complete a task, including the checks needed to judge the result. Access and procedure serve different purposes.
How can you tell whether an AI agent skill generalizes?
Test the skill on unfamiliar tasks kept outside the optimization process, and compare results with the original skill and a no-skill baseline. Include different file names and layouts, plus nearby tasks where the skill should remain unused. Measure both work quality and unnecessary effort.
When should a skill become code or model training instead?
Use code for checks and repeated operations that can be performed reliably and mechanically. Keep changing policies and judgment-based procedures in editable guidance. A broadly useful, frequently repeated capability may justify model training if its benefits outweigh the cost and can be evaluated.
⬅️ AI Builds AI#3: Why AI Infrastructure Giants Are Racing Into Robotics






