Writing · Architecture

The agent memory problem: notes from building one

2026 · 5 min read

The first time I noticed it, I was watching our agent fail at a task it had succeeded at twenty minutes earlier.

Same prompt, same environment, same model. The first run, it navigated through a workflow cleanly. The second run, it stumbled at the same step three times before getting it right. The third run, it stumbled differently. By the end of the afternoon I had a hundred runs, and I could plot the success rate as a flat line that did not move with model upgrades, prompt tuning, or temperature.

The model was not getting worse. It was failing to learn anything.

That was the moment I understood that the agentic AI conversation in 2026 is being held at the wrong layer. We argue about which model. We argue about which framework. We argue about orchestrators and protocols. Almost nobody is asking the prior question: what shape should an agent's memory take?

The pattern of failure

If you have shipped an agent to production, you have probably seen some version of this:

  • It reasons brilliantly inside a single trajectory.
  • It forgets that reasoning the next time the same situation arises.
  • It re-discovers it, sometimes correctly, sometimes not.
  • Token cost climbs as you bolt on longer context windows to compensate.
  • Reliability does not climb with it.

This is not a "scale the model" problem. I tried that. We have all tried that. A bigger model with no memory is just a more expressive amnesiac.

It is also not a "vector store" problem in the way the industry has framed it. Vector retrieval gives an agent the ability to look up text it has seen before. That is useful for question answering. It is almost useless for procedural tasks, because the unit of memory an agent needs is not "a paragraph I once read." It is "in state X, action A led to outcome O, with probability P."

That is a graph. Specifically, a graph with states as nodes, actions as edges, and outcomes as labels you accumulate as the agent acts. The agent's memory has to have the same shape as the agent's job.

What we built

Our agent platform has two memory components and three modes of operation.

The two memory components:

  • Short-term memory (STM): the current trajectory. State, action, observation, outcome, in sequence. Bounded.
  • Long-term memory (LTM): a persistent graph. Nodes are states the agent has been in. Edges are actions it has taken between them. Labels accumulate across runs: how often this action worked, under what conditions, with what side effects.

The three modes:

  • Navigate: use the LTM. If the agent has been here before and knows what worked, just do it. No LLM call.
  • Explore: the LTM does not cover the situation. Try something reasonable, log the outcome, expand the graph.
  • Consult: when even exploration is uncertain, escalate to the LLM. The LLM proposes; the agent verifies and stores the outcome.

The LLM, in this architecture, is not the brain. It is the tongue. It is excellent at language, ambiguity, and creative proposal. It is bad at remembering what you taught it last Tuesday. So we stopped asking it to.

What the numbers said

I am wary of cherry-picking a benchmark, so I will tell you the one I tracked because it surprised me.

We ran a Smart Room prototype: an agent operating a simulated environment with a fixed set of objectives. Before the memory rework, Navigate-mode success was around 2%. Not a typo. The agent had no idea what had worked before, so almost every run was a fresh exploration that re-learned nothing.

After we introduced the STM/LTM graph and the three-mode operation, the numbers over four days looked like this:

  • Navigate-mode success: 2% climbing to roughly 80%
  • LLM calls per task: dropping toward zero by day four
  • Decisions logged: 866
  • Learned associations stored procedurally: 214

The most interesting line is the LLM-calls one. The agent did not need a bigger model. It needed to stop asking the model the same questions over and over. By day four it was running mostly on remembered structure, calling the LLM only at genuine novelty.

That is what cheap, reliable agentic AI looks like when you draw it on a chart.

Why this matters to a CTO

If you are evaluating agentic AI vendors right now, three questions are worth asking before any conversation about the model:

  1. What is the shape of the agent's memory? If the answer is "a vector database," the vendor has not solved the problem. They have stored the symptom.
  2. Can the agent distinguish what it remembers from what it is guessing? Without that distinction, you cannot audit it, and you cannot trust it in production.
  3. Does the cost per task fall over time as the agent learns, or stay flat? If costs are flat, the agent is not learning, it is just performing. You will pay forever for a service that should be getting cheaper.

The Gartner number that has been making the rounds (over 40% of agentic AI projects will be cancelled by the end of 2027, Gartner predicted in June 2025) is plausible to me precisely because most of the projects I have seen evaluated do not pass any of those three tests. They cannot. They were not architected to.

What I would tell my past self

The thing that delayed us was the assumption that the LLM was supposed to be doing the thinking. Once we stopped treating the LLM as the brain and started treating it as a narrow, expensive specialist that we would consult sparingly, the architecture clarified.

Agents are not language models with tools. They are stateful programs that occasionally borrow language from a model.

If you internalize that, the question "which model should we use" becomes far less urgent than "what does our agent know, and how does it know it." That is the question I wish the industry was arguing about. It is the one I expect to define which agentic projects survive 2027 and which become the cancellation statistic.

Drafted in 2026
Updated for site in October 2026