The lesson in one minute
What you'll be able to explain
- The model is stateless; memory is what your code puts back into the prompt.
- Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
- Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
- Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
- Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
- Keep task state in a database; the model proposes, code validates and writes.
Level 1
The practitioner's guide
In one sentence
A model remembers nothing between calls, so "memory" is the system you build around it: what you keep from a conversation, what you store for months, what you put back into the prompt, who may see it, and where the state of a long task lives so that a crash doesn't lose it.
When you need it
Any product where the second conversation should know about the first, or where one conversation runs long enough to outgrow the window: assistants that remember preferences, agents that resume a multi-step job, support bots with long threads, anything multi-tenant. You don't need long-term memory for a one-shot task, and you don't need a database for a job that finishes in one call. The tells: users repeating themselves every session; a chat that gets slower, dearer and vaguer as it goes (this lesson's unmanaged history climbs without end, while the managed one levels off just under its 300-token budget from about turn 17); an agent that "forgot" a fact the user corrected an hour ago; or a job that has to start over because the process died halfway.
Your options
From nothing to a full system, each layer added on top of the previous one:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| The whole history in the prompt | Resends every turn on every call | Nothing is ever forgotten inside one session | Tokens that grow every turn; quality that fades with length | Your prompt |
| Short-term memory with a budget | Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave | The history never exceeds the budget, however long the session | Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call | Your code, per conversation |
| Notes the agent reads on demand | The model writes and reads files or notes outside the window through a tool | Facts and progress survive a reset of the window | Tool calls per read and write, and storage you control | A tool and a directory or table |
| Long-term memory, typed and recalled by similarity | Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few | The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased | An embedding per record, a write policy, a partitioned store, export and deletion paths | A store your code owns |
| Task state in a database | Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions | A run survives a crash and resumes at the first unfinished step; no step can be skipped | A schema and a validator per task type | A database |
How to choose
Decide separately for the conversation, for facts that outlive it, and for the state of a job, because they fail differently.
- A chat that runs long: short-term memory with a budget, always. Keep the recent turns verbatim (the order number the user just typed) and summarize the rest.
- Anything the user would be annoyed to repeat (preferences, corrections, how they like a task done): long-term memory, written selectively. A good assistant's notebook holds "prefers morning meetings", not "said thanks at 3:02pm", and never a password.
- A job with steps that must happen in order and may outlive a process:
task state in a database, with the model proposing and code holding the
pen. In this lesson the model's attempt to mark
matchdone whilefetch_paymentsis still open is refused, and after a crash the new process resumes atmatchfrom the file, not from a conversation that no longer exists. - Many customers on one system: partition the store by tenant and user,
chosen by your authentication layer, never by the query. A query that
quotes another tenant's secret word for word, with
OR 1=1appended, returns nothing from outside the caller's partition here, because there is no path to search it. - Whatever you pick: the model proposes, your code decides what is written. Arrows into the stores never come straight from the model.
What it costs
Short-term memory costs the summarizer (free and crude if extractive; a small model call, with less lost, in production) and the detail it drops. Long-term memory costs an embedding per record at write time, a similarity search per recall, and the engineering around it: a write policy, supersession with provenance, per-user export and deletion. The last two are not optional where privacy law applies; the GDPR's Article 17 gives a person the right to erasure "without undue delay". Task state costs a schema, a validator and a transaction per update, and buys runs that people and monitoring can query. What none of this costs is model quality: memory is retrieval over your own history, and it lives entirely in your code.
What breaks
- Context rot. Quality drops and cost climbs as a session grows. Budget the history, summarize the old turns, promote durable facts to long-term memory, or reset with a written hand-off.
- Remembering everything. Small talk crowds out facts, and a stored secret is read back into every future prompt and every export. Skip chatter, refuse anything that looks like a credential, and strip sensitive data before a note is written.
- Silent overwrites. The user said April, then corrected to July; a store that overwrites cannot explain why the agent ever said April. Supersede under the same key, recall only the current value, and keep the history with its source.
- Isolation by filter. A global index filtered after the search is one bug away from a leak; in this lesson's figure the other tenant's secret scores highest against the adversarial query. Partition first; search inside the partition only; test with adversarial queries.
- Progress kept in the conversation. It is lost on a crash, and it can be summarized, trimmed or misread. Keep it in a database and let code enforce the order of steps.
- Paths that escape. A memory tool that maps names to files must
reject
../and its encodings, or a request for/memories/../secretsreads outside the store.
In the wild
Park et al. (2023), Generative Agents, gave twenty-five
simulated agents a memory stream of natural-language observations,
retrieved by relevance, recency and importance and periodically reflected
into higher-level memories. Packer et al. (2023), MemGPT, treat the window
as fast memory and external storage as slow memory, paging between them
the way an operating system does. Claude's memory tool is the note-taking
option as a product: the model issues view, create, replace, insert,
delete and rename commands against a /memories directory that your
application maps onto storage it controls, its system instruction tells the
model to assume the window may be reset at any moment, and it pairs with
server-side compaction. LangGraph names the same split: thread-scoped
short-term memory held by a checkpointer, long-term memory in a store with
namespaces, and the semantic, episodic and procedural kinds. SQLite's
transactions are the all-or-nothing writes the task store relies on.
Go deeper
Level 2 builds each store in plain Python: a short-term memory whose summary only gets the room the recent turns leave, a long-term memory with three kinds of record recalled by cosine similarity, the write policy, supersession and per-tenant partitions with the leak that a filter would allow drawn as a figure, and a SQLite task store that refuses a skipped step and resumes after a crash. If you only needed to choose, you are done.
Level 2
How it works, from scratch
What follows builds the four pieces, each small enough to read in one sitting, and shows what each one prevents.
A language model remembers nothing between calls. Every call starts from a blank page plus whatever you put in the prompt. "Memory" is therefore entirely your system: what you keep, where you keep it, what you put back into the prompt, and who is allowed to see it. This lesson builds four pieces: short-term memory, long-term memory with three kinds of record, the safety rules around it (write policy, conflicts, isolation, deletion), and task state kept in a database instead of in the conversation.
Figure 1 · Diagram
flowchart LR
subgraph Prompt["What the model sees on this call"]
S[System prompt<br/>+ summary of older turns]
R[Recalled long-term memories]
M[Recent messages, word for word]
end
STM[(Short-term memory<br/>this conversation)] --> S
STM --> M
LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R
DB[(Task state<br/>SQLite)] -->|current step| S
Model[Model] -->|proposes updates| Code[Your code]
Code -->|validated writes| LTM
Code -->|validated writes| DB
Chapter 1
Short-term memory: the conversation, inside a budget
Everyday picture A flip chart in a long meeting. When the page fills up, you don't find a bigger pad. You tear off the old pages and start a fresh one with one line at the top: "Earlier: agreed on the budget, rejected vendor B."
Tiny worked example Six question-and-answer exchanges about invoices, 17 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the first sentence of each older message, but all eight of those would take 65 tokens, so the oldest drop out, one at a time, until the rest fit in 40:
summary : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
Question 4 is about invoices; Answer 4 lists the invoice (36 tokens)
messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ... (verbatim, 70 tokens)
Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 are gone from the prompt entirely. That's the price of a fixed budget, and it's why anything worth keeping forever belongs in long-term memory (section 2), not in the conversation.
Figure 2 · Drawn from the lesson's code
Sending the whole history climbs without end, while the managed history levels off just under its 300-token budget
The code ShortTermMemory.context() returns (summary, messages). The
summary goes in the system prompt rather than as a message, so user and
assistant turns still alternate as the API requires.
In code: ShortTermMemory.add appends a message and
ShortTermMemory.tokens totals the history; first_sentences is the
default summarizer, keeping the first sentence of each folded message.
ShortTermMemory.context gives the summary only the budget the recent messages leave
and drops the oldest folded messages until it fits. A production system
often re-summarizes the summary with an LLM call instead, trading an
extra call for losing less.
Why it matters Context rot, where quality drops as a session grows, is one of the most common agent failures. Summarize, trim, or reset with a written hand-off.
Chapter 2
Long-term memory: three kinds of record
Everyday picture Three notebooks: a diary of what happened (episodic: "on 18 Sept Alice rejected Globex"), an encyclopedia of facts (semantic: "Alice's fiscal year starts in April"), and a habit, the way you've learned to do something (procedural: "to reconcile, match each invoice to its payment, then list mismatches").
Tiny worked example Alice has one memory of each kind. The question "what
happened with Globex?" is embedded and compared with each memory (cosine
similarity, see primer.ml.embeddings.similarity), and the diary entry
comes back first. Asking with kinds=("procedural",) searches only habits.
| kind | stored as | typical recall trigger |
|---|---|---|
| episodic | dated event | "what happened with...", "last time..." |
| semantic | keyed fact (fiscal_year_start) |
any question the fact answers |
| procedural | how-to steps | "how do I...", before starting a known task |
Figure 3 · Drawn from the lesson's code
Each question recalls the right kind: the Globex question scores the diary entry highest, the how-to question the procedure
primer.agents.rag) over the agent's own history.In code: LongTermMemory.remember stores a MemoryRecord of one kind
with its embedding and reports back in a WriteResult.
LongTermMemory.recall returns the current records in a scope most similar
to the query, optionally limited to some kinds.
Chapter 3
The hard parts: what to write, conflicts, isolation, deletion
What to write. Everyday picture: a good assistant's notebook has
"prefers morning meetings", not "said thanks at 3:02pm", and never your
bank PIN. worth_remembering skips small talk and refuses anything that
looks like a secret, since memory is read back into prompts and shown in
exports.
Conflicts. Worked example: Alice said her fiscal year starts in April,
then corrected it to July. Both are stored under the key fiscal_year_start.
The April record is marked superseded by the July one, recall returns only
July, and history() still shows both with their source, so "why did the agent
think April?" has an answer.
Isolation. Everyday picture: separate locked filing cabinets per
company, not one cabinet with a "please only read your own folder" sign.
Multi-tenant means one system serves many customers (tenants). Every
memory call names a Scope (tenant + user), and storage is partitioned
by it, so a lookup starts inside the caller's own cabinet and can't reach
another.
Figure 5 · Diagram
flowchart TD Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol] P --> S[Similarity search<br/>inside this partition only] S --> R[Results] A[(acme/alice partition)] -.-x S
Figure 4 · Drawn from the lesson's code
Acme's secret scores a perfect match to Carol's quoted query, yet it sits outside her partition and is never searched
Deletion. Users may need to see, correct or delete what's remembered,
and privacy law such as the GDPR (the EU's General Data Protection
Regulation) can require it. export(scope) shows everything including
superseded facts, and delete_user(scope) erases one user without touching
colleagues.
In code: LongTermMemory.remember applies the write policy and marks
an older record with the same key as superseded; LongTermMemory.history
lists every value a key has had. LongTermMemory keeps one partition per
Scope, and LongTermMemory.export, LongTermMemory.forget and
LongTermMemory.delete_user show, remove one record, and erase a user.
Chapter 4
Task state outside the model
Everyday picture A checklist on a clipboard. The assistant can suggest ticking a box, but only the supervisor holds the pen. If the assistant goes home sick, the next person picks up the clipboard and starts at the first unticked box.
Tiny worked example
steps : fetch_invoices, fetch_payments, match
model proposes : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"} -> written
model proposes : {"step": "match", "status": "done"} -> refused: the current step is fetch_payments
process crashes after fetch_payments is done
new process, same file : next_step() -> "match"
Figure 6 · Diagram
sequenceDiagram
participant M as Model
participant C as Your code
participant D as SQLite
M->>C: propose {"step": "fetch_invoices", "status": "done"}
C->>D: read current step
D-->>C: fetch_invoices
C->>D: write status=done (transaction)
M->>C: propose {"step": "match", "status": "done"}
C-->>M: refused: current step is fetch_payments
Note over C,D: crash, restart
C->>D: next_step?
D-->>C: match
In code: TaskStateStore keeps tasks and steps in SQLite.
TaskStateStore.create_task writes a task and its steps in one transaction,
TaskStateStore.next_step finds the first unfinished step, and
TaskStateStore.apply validates a proposal, raising InvalidUpdate for a
bad status or a skipped step, before writing it.
Why it matters Authoritative state in a database makes runs inspectable, resumable and auditable. That's much of the difference between a demo and a production system.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: How would you design long-term memory for a multi-tenant agent platform?Think it through, then reveal
A: Every read and write takes a scope (tenant, user) from the authenticated session, and storage is partitioned by it: separate namespaces, collections or row-level security, never a global index filtered after the search. Store typed records (episodic, semantic with keys, procedural) with embeddings, source and time. Apply a write policy (durable, non-secret), supersede conflicting facts instead of overwriting, retrieve the top few by similarity within scope, and support export and deletion per user. Test isolation with adversarial queries.
Question 2Q: A long support chat gets worse and more expensive over time. Why, and what helps?Think it through, then reveal
A: Every call re-sends the whole history, so cost rises, and models use details in the middle of long contexts less reliably. Keep a budget: recent turns verbatim, older ones summarized, important facts promoted to long-term memory, or reset with a written hand-off.
Question 3Q: The user corrects a fact the agent remembered. What should happen?Think it through, then reveal
A: Store the new value under the same key, mark the old one as superseded by the new (don't delete it silently), and recall only current records. Keep the source of each so the change is explainable.
Question 4Q: Why keep task state in a database when the model can "remember" progress in the conversation?Think it through, then reveal
A: The conversation is lost on a crash, and it can be summarized, trimmed or misread. A database is authoritative, survives restarts, can be queried by people and monitoring, and lets code enforce rules (no skipping steps) that a prompt can only request.
Primary sources
The papers behind this lesson
It introduced a memory stream of observations retrieved by relevance, recency and importance, plus periodic reflection into higher-level memories, a template for long-term agent memory.
The paper ↗It treats the context window like RAM and external storage like disk, with the agent paging information in and out, which is the short-term/long-term split made explicit.
The paper ↗Researcher's shelf
Further reading
- Lilian Weng, LLM Powered Autonomous Agents (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/
- LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/
- SQLite, transactions: https://www.sqlite.org/lang_transaction.html
- GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.