Agent Memory

How agents remember: the context window as working memory, summarization and compaction, and long-term stores with retrieval.

What is it?

LLMs have no memory of their own between API calls. Everything an agent 'remembers' is something your code chose to put into the next request. Agent memory is the set of techniques for deciding what to keep, what to compress, what to store outside the model, and what to bring back when it is needed.

It helps to borrow terms from human memory:

  • Working memory (short-term): the current context window - the system prompt, the message history, tool calls and results of this run. It is fast and precise, but limited in size, and every token in it is paid for on every call.
  • Long-term memory: information stored outside the model - files, a database, a vector store - that survives across sessions and can be far larger than any context window.
  • Long-term memory is often split further: semantic memory (facts: 'the user prefers Python'), episodic memory (what happened: 'last Tuesday we migrated the billing DB and it failed on step 3'), and procedural memory (how to do things: instructions and learned rules, often kept in the system prompt or a rules file).

Working memory management. A long agent run fills the context with old tool outputs. Problems appear before you hit the hard limit: cost and latency grow every step, and models attend less reliably to details buried in a very long context (sometimes called context rot). Techniques, from simplest to most involved:

  • Trim tool outputs at the source: return 20 relevant lines, not 2,000.
  • Sliding window: keep only the last N turns. Simple, but the agent forgets the original goal unless you always keep the first user message.
  • Clear old tool results: replace bulky results from earlier steps with a short placeholder like '[result removed - re-run read_file if needed]', keeping the tool_use/tool_result structure valid.
  • Summarization / compaction: when the history passes a threshold, ask the model to summarize the older part (decisions made, facts learned, open questions, files touched), then continue with summary + recent turns. This keeps the essentials at a fraction of the tokens.
  • Scratchpad / notes file: give the agent a tool to write its own progress notes (a to-do list, findings) to a file, and to read them back. Its plan then survives even aggressive compaction.

Long-term memory usually works through tools or automatic retrieval:

  • Explicit memory tools: remember(fact) and recall(query). The model decides what is worth saving. Simple and transparent; you can show users what was stored and let them delete it.
  • Retrieval-based memory: store memories as embeddings (vectors that capture meaning, see Embeddings) in a vector store, and before each turn retrieve the few memories most similar to the current message and insert them into the prompt. This is RAG applied to the agent's own past.
  • Structured profiles: for well-known fields (name, timezone, plan tier), a normal database row beats fuzzy retrieval.

Memory needs housekeeping: memories go stale ('user works at Acme' after they changed jobs), conflict, or contain sensitive data. Store a timestamp and source with each memory, prefer newer facts, let users view and delete what is stored, and never save secrets such as passwords or API keys.

Some model providers, including Anthropic, offer built-in context-management features (for example automatic compaction or clearing old tool results, and a client-side memory tool). They implement the same ideas; check the current documentation before relying on a specific one.

Explain like I'm 10

Working memory is your desk: only so many papers fit, but everything on it is right in front of you. Compaction is clearing the desk at the end of a meeting by writing a one-page summary and filing the rest. Long-term memory is the filing cabinet in the corner: huge, but you only benefit from it if you know what to pull out when you need it - and if someone throws away the outdated folders.

Examples

Compaction: summarize old turns when over budget (runnable)

// Token counting approximated by words, just for the demo.
const countTokens = (msgs) => msgs.reduce((n, m) => n + m.content.split(" ").length, 0);

// Fake summarizer LLM: keeps only sentences containing decisions or preferences.
function summarize(msgs) {
  const keep = msgs.filter((m) => ["decided", "prefers", "deadline"].some((w) => m.content.includes(w)));
  return "Summary of earlier conversation: " + keep.map((m) => m.content).join(" ");
}

function compact(system, msgs, budget, keepRecent) {
  if (countTokens(msgs) <= budget) return { system, msgs };
  let cut = msgs.length - keepRecent;
  while (cut < msgs.length && msgs[cut].role !== "user") cut++;   // recent part must start with a user turn
  const summary = summarize(msgs.slice(0, cut));
  return { system: system + " " + summary, msgs: msgs.slice(cut) };
}

let system = "You are a project assistant.";
let msgs = [
  { role: "user", content: "Let's plan the website relaunch for the bakery client" },
  { role: "assistant", content: "Sure. What matters most to them?" },
  { role: "user", content: "The client prefers a green colour scheme and hates pop ups" },
  { role: "assistant", content: "Noted. Any hard dates?" },
  { role: "user", content: "The deadline is the 3rd of June because of their anniversary sale" },
  { role: "assistant", content: "Then we should freeze the design two weeks before." },
  { role: "user", content: "Agreed, we decided to use the existing logo, no redesign" },
  { role: "assistant", content: "Great, that saves a week of work on branding and approvals." },
  { role: "user", content: "Now draft the homepage sections please" },
];

console.log("before:", msgs.length, "messages,", countTokens(msgs), "tokens");
const out = compact(system, msgs, 60, 3);
console.log("after :", out.msgs.length, "messages,", countTokens(out.msgs), "tokens");
console.log("system now:", out.system);
console.log("first kept message:", out.msgs[0].role, "-", out.msgs[0].content);

Old turns are folded into a short summary in the system prompt; recent turns stay verbatim. The cut point is moved so the kept history starts with a user message, keeping roles valid. A real summarizer is an LLM call with instructions on what to preserve.

Retrieval-based long-term memory with tiny vectors (runnable)

const STOP = new Set(["the", "is", "a", "for", "of", "should", "be", "how", "which", "what", "my", "to", "on", "does"]);
function embed(text) {                  // toy bag-of-words "embedding"
  const v = {};
  for (const w of text.toLowerCase().split(/[^a-z0-9]+/)) if (w && !STOP.has(w)) v[w] = (v[w] || 0) + 1;
  return v;
}
function cosine(a, b) {
  let dot = 0, na = 0, nb = 0;
  for (const k in a) { na += a[k] * a[k]; if (b[k]) dot += a[k] * b[k]; }
  for (const k in b) nb += b[k] * b[k];
  return na && nb ? dot / Math.sqrt(na * nb) : 0;
}

const memory = [];
function remember(text, day) { memory.push({ text, day, vec: embed(text) }); }
function recall(query, k) {
  const q = embed(query);
  return memory
    .map((m) => ({ ...m, score: cosine(q, m.vec) }))
    .filter((m) => m.score > 0)
    .sort((a, b) => b.score - a.score)
    .slice(0, k);
}

remember("User prefers Python for data scripts", 1);
remember("User dog is named Biscuit", 2);
remember("Project deadline is 3 June", 3);
remember("User dislikes long answers, keep replies short", 4);
remember("Project deadline moved to 10 June", 9);

for (const q of ["which language for data scripts", "how long should replies be", "project deadline"]) {
  const hits = recall(q, 2);
  console.log(q, "->", hits.map((h) => h.text + " (" + h.score.toFixed(2) + ", day " + h.day + ")").join(" | "));
}
// Pure similarity ranked the STALE deadline first (shorter text, higher score).
// Fix: among relevant hits, prefer the newest, and label memories with their date.
const relevant = recall("project deadline", 3).filter((m) => m.score >= 0.5);
const newest = relevant.sort((a, b) => b.day - a.day)[0];
console.log("prompt context: - (saved day " + newest.day + ") " + newest.text);

Each memory is stored with a vector; at question time the most similar memories are retrieved and added to the prompt. The conflicting deadlines show a real trap: similarity alone ranks the outdated deadline first, so you need timestamps and a recency rule (here: newest relevant memory wins). Real systems use a real embedding model and a vector store.

Long-term memory as tools, persisted in a JSON file

import json, time
from pathlib import Path

MEMORY_FILE = Path("memory.json")

MEMORY_TOOLS = [
    {"name": "remember",
     "description": "Save a durable fact about the user or project for future sessions "
                    "(preferences, decisions, deadlines). Do NOT save secrets, passwords or "
                    "one-off details. Write the fact as one self-contained sentence.",
     "input_schema": {"type": "object",
                      "properties": {"fact": {"type": "string", "maxLength": 300}},
                      "required": ["fact"]}},
    {"name": "recall",
     "description": "Search saved facts by keyword. Use at the start of a task when past "
                    "preferences or decisions could matter. Returns newest matches first.",
     "input_schema": {"type": "object",
                      "properties": {"query": {"type": "string"}},
                      "required": ["query"]}},
]

def _load() -> list[dict]:
    return json.loads(MEMORY_FILE.read_text()) if MEMORY_FILE.exists() else []

def remember(fact: str) -> str:
    items = _load()
    items.append({"fact": fact.strip(), "saved_at": time.strftime("%Y-%m-%d %H:%M")})
    MEMORY_FILE.write_text(json.dumps(items, indent=2))
    return "remembered"

def recall(query: str) -> str:
    words = [w for w in query.lower().split() if len(w) > 2]
    hits = [m for m in _load() if any(w in m["fact"].lower() for w in words)]
    hits.sort(key=lambda m: m["saved_at"], reverse=True)
    return "\n".join(f"[{m['saved_at']}] {m['fact']}" for m in hits[:10]) or "no saved facts match"

# Add MEMORY_TOOLS to your agent's tool list and route "remember"/"recall" in run_tool.
# Upgrade path: swap keyword matching for embeddings + a vector store (e.g. Chroma).

The model decides what to save and when to look things up. Because memory lives in a plain file, users can inspect and edit it - an important trust and privacy feature.

Compacting an agent's history with an LLM summary

import json
import anthropic

client = anthropic.Anthropic()

def render(m: dict) -> str:
    """Turn one message (string or content blocks) into readable text for the summarizer."""
    if isinstance(m["content"], str):
        return f"{m['role']}: {m['content']}"
    parts = []
    for b in m["content"]:
        b = b if isinstance(b, dict) else b.model_dump()      # SDK blocks are pydantic models
        if b["type"] == "text":
            parts.append(b["text"])
        elif b["type"] == "tool_use":
            parts.append(f"[called {b['name']} {json.dumps(b['input'])}]")
        elif b["type"] == "tool_result":
            parts.append(f"[result: {str(b['content'])[:500]}]")
    return f"{m['role']}: " + " ".join(parts)

def compact(messages: list, keep_recent: int = 6) -> list:
    # Cut only at a plain user message, so no tool_use is separated from its tool_result.
    cut = len(messages) - keep_recent
    while cut < len(messages) and not (messages[cut]["role"] == "user"
                                       and isinstance(messages[cut]["content"], str)):
        cut += 1
    if cut <= 1 or cut >= len(messages):
        return messages                                         # nothing safe to compact
    transcript = "\n".join(render(m) for m in messages[:cut])
    summary = client.messages.create(
        model="claude-opus-5-5", max_tokens=2000,
        system="You compress agent transcripts. Keep: the original goal, decisions, facts "
               "learned, files touched, errors seen, and open next steps. Be terse.",
        messages=[{"role": "user", "content": transcript}],
    ).content[0].text
    first = messages[cut]
    return [{"role": "user",
             "content": f"Summary of earlier work:\n{summary}\n\nContinue with: {first['content']}"}] \
           + messages[cut + 1:]

# In the loop: if resp.usage.input_tokens > THRESHOLD: messages = compact(messages)

The tricky part is choosing a safe cut point: never separate a tool_use from its tool_result. Folding the summary into the first kept user message keeps roles alternating.

How it works

Every call is built fresh: system prompt (instructions + possibly retrieved long-term memories) + message history (working memory, possibly compacted) + the new user message. Memory management is simply the code that assembles that request.

Compaction trades detail for space. A good summary prompt lists exactly what must survive: the goal, constraints, decisions, facts learned, what was tried and failed, and the next steps. Without that, summaries drop the one detail that mattered.

Retrieval-based memory follows the RAG pipeline: write (embed + store with metadata such as time, user and source), read (embed the current query, find nearest neighbours, filter by user and recency), and inject (add the top few into the prompt, clearly labelled as memories that may be outdated). Scoping by user id is essential - one user's memories must never leak into another's prompt.

   each model call is assembled from:
  +------------------------------------------+
  | system: instructions                     |
  |         + recalled long-term memories <--+-- vector store /
  |         + summary of old turns           |   files / DB
  | messages: recent turns (verbatim)        |
  |           + new user message             |
  +------------------------------------------+
        |  context full?  -> compact old turns
        |  new durable fact? -> remember() -> store

Why does it exist?

Context windows are finite and every token costs money and time on every step. At the same time, useful assistants must stay coherent through long tasks and remember users across sessions. Memory techniques let an agent keep the important parts of a long history and a large body of past knowledge without stuffing everything into every request.

When to use it

Use compaction or result-clearing for long-running agents (coding sessions, research tasks) whose history grows past a comfortable size. Use long-term memory for assistants that serve the same user repeatedly and benefit from preferences and past decisions. Use a notes/scratchpad tool for multi-hour tasks that need a persistent plan.

When not to use it

Short, single-session tasks need no memory system - the message history is enough. Do not add vector-based memory when a few structured fields in a database would do. Avoid persistent memory where privacy rules forbid storing user data, or where stale memories could cause harm (medical, legal) without careful review.

Common mistakes

  • Assuming the model remembers previous API calls - it does not; you must resend or re-inject.

  • Cutting history in the middle of a tool_use/tool_result pair, causing API errors.

  • Using a sliding window that drops the original task, so the agent forgets what it was doing.

  • Summaries that lose the decisive details because the summary prompt did not say what to keep.

  • Storing every message as a memory, so retrieval returns noise.

  • Never expiring or updating memories, so outdated facts override current ones.

  • Mixing memories across users or tenants in one unfiltered vector index.

  • Saving secrets or sensitive personal data into long-term memory.

Practice exercises

  1. Easy:

    In the compaction demo, change the budget to 200 and to 30. What changes, and why?

  2. Easy:

    Add a memory 'User prefers TypeScript for web scripts' on day 12 to the retrieval demo and query 'which language for scripts'. Explain the ranking.

  3. Medium:

    Write a summary prompt for a coding agent's compaction step. List the exact sections it must output and test it on a transcript of one of your own agent runs.

  4. Medium:

    Add remember and recall to mini_agent.py from Build an Agent from Scratch, and verify a fact saved in one run is recalled in the next.

  5. Hard:

    Replace the JSON memory with Chroma and sentence-transformers: store each fact with user_id and timestamp metadata, retrieve with a where-filter on user_id, and break ties by recency.

Interview questions

What is working memory for an LLM agent?

The contents of the current context window: system prompt, message history, tool calls and results. It is the only thing the model can see on a call, and it is limited and paid for on every request.

How does compaction work and what are its risks?

When history grows too large, older turns are summarized (usually by the model) and replaced with the summary while recent turns stay verbatim. Risks: losing critical details, breaking tool_use/tool_result pairs at the cut point, and summarizing errors as facts. Mitigate with an explicit summary template and safe cut points.

Explain retrieval-based memory.

Memories are embedded and stored with metadata in a vector store. On each new request, the query is embedded, the most similar memories for that user are retrieved, filtered for recency, and inserted into the prompt. It scales memory far beyond the context window.

Semantic vs episodic vs procedural memory?

Semantic: facts and preferences. Episodic: records of past events and interactions. Procedural: how to do things - instructions, rules and learned workflows, often kept in system prompts or rule files.

How do you handle conflicting or stale memories?

Store timestamps and sources, prefer newer facts, allow updates and deletes (including by the user), periodically consolidate duplicates, and label injected memories as possibly outdated so the model checks them when it matters.

Why not just use a model with a huge context window?

Cost and latency scale with tokens on every call, and models use details in very long contexts less reliably. Long-term memory also must survive across sessions, which a context window cannot do on its own.