rumblr Work in progressWIP

● The AI Primer · Lesson 46 · Part 2: building systems people rely on

Memory and state

You'll be able to explain Short- and long-term memory, tenant isolation, forgetting

Members · open during launch 18 min6 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. The model is stateless; memory is what your code puts back into the prompt.
  2. Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
  3. Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
  4. Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
  5. Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
  6. Keep task state in a database; the model proposes, code validates and writes.

Level 1

The practitioner's guide

In one sentence

A model remembers nothing between calls, so "memory" is the system you build around it: what you keep from a conversation, what you store for months, what you put back into the prompt, who may see it, and where the state of a long task lives so that a crash doesn't lose it.

When you need it

Any product where the second conversation should know about the first, or where one conversation runs long enough to outgrow the window: assistants that remember preferences, agents that resume a multi-step job, support bots with long threads, anything multi-tenant. You don't need long-term memory for a one-shot task, and you don't need a database for a job that finishes in one call. The tells: users repeating themselves every session; a chat that gets slower, dearer and vaguer as it goes (this lesson's unmanaged history climbs without end, while the managed one levels off just under its 300-token budget from about turn 17); an agent that "forgot" a fact the user corrected an hour ago; or a job that has to start over because the process died halfway.

Your options

From nothing to a full system, each layer added on top of the previous one:

Option What it does What it guarantees What it costs Where it lives
The whole history in the prompt Resends every turn on every call Nothing is ever forgotten inside one session Tokens that grow every turn; quality that fades with length Your prompt
Short-term memory with a budget Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave The history never exceeds the budget, however long the session Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call Your code, per conversation
Notes the agent reads on demand The model writes and reads files or notes outside the window through a tool Facts and progress survive a reset of the window Tool calls per read and write, and storage you control A tool and a directory or table
Long-term memory, typed and recalled by similarity Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased An embedding per record, a write policy, a partitioned store, export and deletion paths A store your code owns
Task state in a database Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions A run survives a crash and resumes at the first unfinished step; no step can be skipped A schema and a validator per task type A database

How to choose

Decide separately for the conversation, for facts that outlive it, and for the state of a job, because they fail differently.

  • A chat that runs long: short-term memory with a budget, always. Keep the recent turns verbatim (the order number the user just typed) and summarize the rest.
  • Anything the user would be annoyed to repeat (preferences, corrections, how they like a task done): long-term memory, written selectively. A good assistant's notebook holds "prefers morning meetings", not "said thanks at 3:02pm", and never a password.
  • A job with steps that must happen in order and may outlive a process: task state in a database, with the model proposing and code holding the pen. In this lesson the model's attempt to mark match done while fetch_payments is still open is refused, and after a crash the new process resumes at match from the file, not from a conversation that no longer exists.
  • Many customers on one system: partition the store by tenant and user, chosen by your authentication layer, never by the query. A query that quotes another tenant's secret word for word, with OR 1=1 appended, returns nothing from outside the caller's partition here, because there is no path to search it.
  • Whatever you pick: the model proposes, your code decides what is written. Arrows into the stores never come straight from the model.

What it costs

Short-term memory costs the summarizer (free and crude if extractive; a small model call, with less lost, in production) and the detail it drops. Long-term memory costs an embedding per record at write time, a similarity search per recall, and the engineering around it: a write policy, supersession with provenance, per-user export and deletion. The last two are not optional where privacy law applies; the GDPR's Article 17 gives a person the right to erasure "without undue delay". Task state costs a schema, a validator and a transaction per update, and buys runs that people and monitoring can query. What none of this costs is model quality: memory is retrieval over your own history, and it lives entirely in your code.

What breaks

  • Context rot. Quality drops and cost climbs as a session grows. Budget the history, summarize the old turns, promote durable facts to long-term memory, or reset with a written hand-off.
  • Remembering everything. Small talk crowds out facts, and a stored secret is read back into every future prompt and every export. Skip chatter, refuse anything that looks like a credential, and strip sensitive data before a note is written.
  • Silent overwrites. The user said April, then corrected to July; a store that overwrites cannot explain why the agent ever said April. Supersede under the same key, recall only the current value, and keep the history with its source.
  • Isolation by filter. A global index filtered after the search is one bug away from a leak; in this lesson's figure the other tenant's secret scores highest against the adversarial query. Partition first; search inside the partition only; test with adversarial queries.
  • Progress kept in the conversation. It is lost on a crash, and it can be summarized, trimmed or misread. Keep it in a database and let code enforce the order of steps.
  • Paths that escape. A memory tool that maps names to files must reject ../ and its encodings, or a request for /memories/../secrets reads outside the store.

In the wild

Park et al. (2023), Generative Agents, gave twenty-five simulated agents a memory stream of natural-language observations, retrieved by relevance, recency and importance and periodically reflected into higher-level memories. Packer et al. (2023), MemGPT, treat the window as fast memory and external storage as slow memory, paging between them the way an operating system does. Claude's memory tool is the note-taking option as a product: the model issues view, create, replace, insert, delete and rename commands against a /memories directory that your application maps onto storage it controls, its system instruction tells the model to assume the window may be reset at any moment, and it pairs with server-side compaction. LangGraph names the same split: thread-scoped short-term memory held by a checkpointer, long-term memory in a store with namespaces, and the semantic, episodic and procedural kinds. SQLite's transactions are the all-or-nothing writes the task store relies on.

Go deeper

Level 2 builds each store in plain Python: a short-term memory whose summary only gets the room the recent turns leave, a long-term memory with three kinds of record recalled by cosine similarity, the write policy, supersession and per-tenant partitions with the leak that a filter would allow drawn as a figure, and a SQLite task store that refuses a skipped step and resumes after a crash. If you only needed to choose, you are done.

Level 2

How it works, from scratch

What follows builds the four pieces, each small enough to read in one sitting, and shows what each one prevents.

A language model remembers nothing between calls. Every call starts from a blank page plus whatever you put in the prompt. "Memory" is therefore entirely your system: what you keep, where you keep it, what you put back into the prompt, and who is allowed to see it. This lesson builds four pieces: short-term memory, long-term memory with three kinds of record, the safety rules around it (write policy, conflicts, isolation, deletion), and task state kept in a database instead of in the conversation.

Figure 1 · Diagram

Reading it: the box on the left is the only thing the model ever sees, and it's rebuilt from scratch on every call. The three stores on the right feed it. Arrows into the stores come from your code, never directly from the model. The model proposes and your code decides what gets written.

Chapter 1

Short-term memory: the conversation, inside a budget

Everyday picture A flip chart in a long meeting. When the page fills up, you don't find a bigger pad. You tear off the old pages and start a fresh one with one line at the top: "Earlier: agreed on the budget, rejected vendor B."

Tiny worked example Six question-and-answer exchanges about invoices, 17 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the first sentence of each older message, but all eight of those would take 65 tokens, so the oldest drop out, one at a time, until the rest fit in 40:

summary  : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
           Question 4 is about invoices; Answer 4 lists the invoice              (36 tokens)
messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ...   (verbatim, 70 tokens)

Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 are gone from the prompt entirely. That's the price of a fixed budget, and it's why anything worth keeping forever belongs in long-term memory (section 2), not in the conversation.

Figure 2 · Drawn from the lesson's code

0 5 10 15 20 25 30 35 40 turn 0 200 400 600 800 1000 1200 1400 tokens of history sent Short-term memory under a budget send the whole history recent turns + first-sentence summary budget before summarizing (300)

Sending the whole history climbs without end, while the managed history levels off just under its 300-token budget

Reading it: the x-axis is the turn number in a long conversation and the y-axis is how many tokens of history are sent on that turn. Unmanaged, the line climbs forever, and so do cost and latency. The model also gets worse at using details buried in the middle. Managed, it climbs until the history first passes the 300-token budget (turn 9), drops as the older turns fold into a summary, then climbs back and stays flat just under 300 from about turn 17 on. It stays flat because the summary only gets the room the recent messages leave: each new turn folds one more exchange in, and the oldest one drops out of the summary to make space.

The code ShortTermMemory.context() returns (summary, messages). The summary goes in the system prompt rather than as a message, so user and assistant turns still alternate as the API requires.

In code: ShortTermMemory.add appends a message and ShortTermMemory.tokens totals the history; first_sentences is the default summarizer, keeping the first sentence of each folded message. ShortTermMemory.context gives the summary only the budget the recent messages leave and drops the oldest folded messages until it fits. A production system often re-summarizes the summary with an LLM call instead, trading an extra call for losing less.

Why it matters Context rot, where quality drops as a session grows, is one of the most common agent failures. Summarize, trim, or reset with a written hand-off.

Chapter 2

Long-term memory: three kinds of record

Everyday picture Three notebooks: a diary of what happened (episodic: "on 18 Sept Alice rejected Globex"), an encyclopedia of facts (semantic: "Alice's fiscal year starts in April"), and a habit, the way you've learned to do something (procedural: "to reconcile, match each invoice to its payment, then list mismatches").

Tiny worked example Alice has one memory of each kind. The question "what happened with Globex?" is embedded and compared with each memory (cosine similarity, see primer.ml.embeddings.similarity), and the diary entry comes back first. Asking with kinds=("procedural",) searches only habits.

kind stored as typical recall trigger
episodic dated event "what happened with...", "last time..."
semantic keyed fact (fiscal_year_start) any question the fact answers
procedural how-to steps "how do I...", before starting a known task

Figure 3 · Drawn from the lesson's code

what happened with Globex? how do I reconcile invoices? 0.0 0.2 0.4 0.6 0.8 cosine similarity Recall: which of Alice's memories each question finds memory kind semantic episodic procedural

Each question recalls the right kind: the Globex question scores the diary entry highest, the how-to question the procedure

Reading it: each group of bars is one question, and each bar is one of Alice's memories, coloured by kind. The tallest bar in each group is what gets recalled. "What happened with Globex?" lights up the diary entry, because only it mentions Globex. "How do I reconcile invoices?" lights up the procedure. This is RAG (retrieval-augmented generation, see primer.agents.rag) over the agent's own history.

In code: LongTermMemory.remember stores a MemoryRecord of one kind with its embedding and reports back in a WriteResult. LongTermMemory.recall returns the current records in a scope most similar to the query, optionally limited to some kinds.

Chapter 3

The hard parts: what to write, conflicts, isolation, deletion

What to write. Everyday picture: a good assistant's notebook has "prefers morning meetings", not "said thanks at 3:02pm", and never your bank PIN. worth_remembering skips small talk and refuses anything that looks like a secret, since memory is read back into prompts and shown in exports.

Conflicts. Worked example: Alice said her fiscal year starts in April, then corrected it to July. Both are stored under the key fiscal_year_start. The April record is marked superseded by the July one, recall returns only July, and history() still shows both with their source, so "why did the agent think April?" has an answer.

Isolation. Everyday picture: separate locked filing cabinets per company, not one cabinet with a "please only read your own folder" sign. Multi-tenant means one system serves many customers (tenants). Every memory call names a Scope (tenant + user), and storage is partitioned by it, so a lookup starts inside the caller's own cabinet and can't reach another.

Figure 5 · Diagram

Reading it: the query never runs against the whole store. It first opens exactly one partition, chosen from the scope that your authentication layer supplies, never from the query text. The crossed dotted line is the point: there is no path from Carol's search to Acme's records, so no clever query can create one.

Figure 4 · Drawn from the lesson's code

−0.2 0.0 0.2 0.4 0.6 0.8 1.0 similarity to Carol's query acme/alice: Alice's fiscal year starts in April.... acme/alice: On 2026-09-18 Alice rejected the vendo... acme/alice: To reconcile invoices, match each vend... acme/alice: Acme plans to acquire Initech in Q4.... globex/carol: Globex prefers invoices in euros.... Carol (globex) quotes Acme's secret: only her partition is searched

Acme's secret scores a perfect match to Carol's quoted query, yet it sits outside her partition and is never searched

Reading it: Carol (tenant Globex) asks a question that quotes Acme's confidential memory word for word. Each bar is that question's similarity to one memory in the whole system. Acme's record scores highest, so a store that searched globally and filtered afterwards is one bug away from leaking it. Hatched bars are outside Carol's partition and are never scored in the real code path.

Deletion. Users may need to see, correct or delete what's remembered, and privacy law such as the GDPR (the EU's General Data Protection Regulation) can require it. export(scope) shows everything including superseded facts, and delete_user(scope) erases one user without touching colleagues.

In code: LongTermMemory.remember applies the write policy and marks an older record with the same key as superseded; LongTermMemory.history lists every value a key has had. LongTermMemory keeps one partition per Scope, and LongTermMemory.export, LongTermMemory.forget and LongTermMemory.delete_user show, remove one record, and erase a user.

Chapter 4

Task state outside the model

Everyday picture A checklist on a clipboard. The assistant can suggest ticking a box, but only the supervisor holds the pen. If the assistant goes home sick, the next person picks up the clipboard and starts at the first unticked box.

Tiny worked example

steps                  : fetch_invoices, fetch_payments, match
model proposes         : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"}   -> written
model proposes         : {"step": "match", "status": "done"}  -> refused: the current step is fetch_payments
process crashes after fetch_payments is done
new process, same file : next_step() -> "match"

Figure 6 · Diagram

Reading it: the model never talks to the database. Every proposal goes through your code, which checks it against the stored truth before writing it inside a transaction (all or nothing). After the crash, the new process learns where to resume from the database, not from a conversation that no longer exists.

In code: TaskStateStore keeps tasks and steps in SQLite. TaskStateStore.create_task writes a task and its steps in one transaction, TaskStateStore.next_step finds the first unfinished step, and TaskStateStore.apply validates a proposal, raising InvalidUpdate for a bad status or a skipped step, before writing it.

Why it matters Authoritative state in a database makes runs inspectable, resumable and auditable. That's much of the difference between a demo and a production system.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: How would you design long-term memory for a multi-tenant agent platform?Think it through, then reveal

A: Every read and write takes a scope (tenant, user) from the authenticated session, and storage is partitioned by it: separate namespaces, collections or row-level security, never a global index filtered after the search. Store typed records (episodic, semantic with keys, procedural) with embeddings, source and time. Apply a write policy (durable, non-secret), supersede conflicting facts instead of overwriting, retrieve the top few by similarity within scope, and support export and deletion per user. Test isolation with adversarial queries.

Question 2Q: A long support chat gets worse and more expensive over time. Why, and what helps?Think it through, then reveal

A: Every call re-sends the whole history, so cost rises, and models use details in the middle of long contexts less reliably. Keep a budget: recent turns verbatim, older ones summarized, important facts promoted to long-term memory, or reset with a written hand-off.

Question 3Q: The user corrects a fact the agent remembered. What should happen?Think it through, then reveal

A: Store the new value under the same key, mark the old one as superseded by the new (don't delete it silently), and recall only current records. Keep the source of each so the change is explainable.

Question 4Q: Why keep task state in a database when the model can "remember" progress in the conversation?Think it through, then reveal

A: The conversation is lost on a crash, and it can be summarized, trimmed or misread. A database is authoritative, survives restarts, can be queried by people and monitoring, and lets code enforce rules (no skipping steps) that a prompt can only request.

Primary sources

The papers behind this lesson

Park et al., Generative Agents: Interactive Simulacra of Human Behavior (2023).

It introduced a memory stream of observations retrieved by relevance, recency and importance, plus periodic reflection into higher-level memories, a template for long-term agent memory.

The paper ↗
Packer et al., MemGPT: Towards LLMs as Operating Systems (2023).

It treats the context window like RAM and external storage like disk, with the agent paging information in and out, which is the short-term/long-term split made explicit.

The paper ↗

Researcher's shelf

Further reading

  • Lilian Weng, LLM Powered Autonomous Agents (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/
  • LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/
  • SQLite, transactions: https://www.sqlite.org/lang_transaction.html
  • GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.