An agent without cross-session persistence starts each session from the context supplied to it. Persistence can reduce repeated briefing, but it also creates privacy, accuracy, isolation, and deletion obligations.
This is the memory problem. Context windows handle the current conversation. Long-term memory — across sessions, across days, across months — needs its own architecture. And it’s harder than it looks.
This article presents a reference design to test, not a certified implementation. Product teams must verify access controls, correction, retention, deletion, recovery, and retrieval quality in their own stack.
What “memory” means
A simplistic view: memory = “the model remembers things between conversations.” Reality is more nuanced. Cognitive science distinguishes types of memory, and AI agent memory benefits from a similar distinction:
Working memory. The current conversation. Held in the context window. Lost when the conversation ends (unless persisted).
Episodic memory. Specific past events. “Last Tuesday, we discussed X.” “Three months ago, you decided Y.”
Semantic memory. General facts. “Your name is Alice.” “You prefer concise responses.” “Your company is in Tallinn.”
Procedural memory. How to do things. “When the user asks for a meeting, use this template.” “When the customer is in tier X, follow process Y.”
Different memory types serve different functions. Use only the layers the product can justify and operate; storing more is not automatically better.
What memory should accomplish
Before architecture, the goals:
Continuity. The agent picks up where it left off. No re-introducing yourself every session.
Personalization. The agent applies your preferences without being asked. Writes in your voice, uses your tools, references your team.
Context preservation. Decisions from past conversations inform current ones. “We decided X last month” should be remembered.
Confirmed preference reuse. The system can reapply an explicit or verified preference in the scopes where it is valid. Repetition alone does not prove that a tool, language, or behavior should become a default.
Privacy and forgetting. What’s remembered, what’s not, what gets deleted. Both for user trust and legal compliance.
These goals can conflict. Continuity may benefit from selected persistence, while privacy and accuracy favor minimization, purpose limits, correction, and deletion. The architecture must make those trade-offs explicit.
The architecture
A typical layered architecture:
┌─────────────────────────────────────┐
│ Working memory (in-context) │ Current conversation
├─────────────────────────────────────┤
│ Session memory (recent) │ Last N conversations
├─────────────────────────────────────┤
│ Episodic memory (long-term) │ Specific past events
├─────────────────────────────────────┤
│ Semantic memory (facts) │ Stable user facts
├─────────────────────────────────────┤
│ Procedural memory (preferences) │ How to behave for this user
└─────────────────────────────────────┘
Each layer needs explicit storage, retrieval, access, provenance, correction, retention, and deletion behavior; multiple logical layers may share a physical store.
We’ll go through each.
Layer 1: Working memory
Already covered in Context engineering. The current conversation in context. For multi-turn conversations, tiered context with recent turns verbatim and older turns summarized.
A hand-off to long-term memory may happen at an explicit checkpoint, a durable event, or session end. Because sessions can end abruptly, persist only approved candidates and make write status observable rather than assuming an end-of-session hook always runs.
Layer 2: Session memory
Recent-session summaries can be kept in limited detail when the product has a justified purpose. Any count, such as the last ten conversations, is an illustrative policy input rather than a default.
Implementation: a summary per session, stored with timestamp and topic. When the user returns, the agent has a quick reference for what’s been happening recently.
{
"session_id": "abc-123",
"user_id": "alice",
"started": "2026-05-14T10:30:00Z",
"ended": "2026-05-14T10:45:00Z",
"topic": "Drafting proposal for Acme Corp",
"summary": "Drafted v1 of the Acme proposal. Decided to lead with the cost-savings angle. Alice will review and send Friday.",
"facts_learned": ["Acme is a current customer", "Alice's deadline is Friday"],
"open_items": ["Alice to review v1 by Thursday"]
}
On a new session, an authorized subset of recent summaries may be loaded after retrieval-quality, relevance, token-budget, and privacy tests. Do not auto-load a fixed number into every context by default.
This is one relatively simple form of cross-session memory, but it still requires isolation, provenance, correction, lifecycle, and retrieval tests.
Layer 3: Episodic memory
Specific past events worth remembering longer-term. Decisions, milestones, important conversations.
These are extracted from sessions when they’re notable. Stored with rich metadata.
{
"event_id": "ev-456",
"user_id": "alice",
"date": "2026-04-22",
"type": "decision",
"description": "Alice decided to migrate from Postgres to ClickHouse for the analytics workload, citing query performance.",
"context_summary": "After 3 weeks of evaluation including performance tests and cost analysis.",
"related_topics": ["infrastructure", "analytics", "database"],
"importance": "high"
}
Retrieval: when relevant to the current conversation, the agent fetches related episodes. Via semantic search (embed the current query, find matching episodes), via topic matching, or via temporal queries (“what happened last month?”).
The challenge is deciding what qualifies as an episode worth remembering. One candidate pattern is for a model to propose decisions, commitments, or milestones at an approved checkpoint, with source spans and confirmation rules. A model-generated importance label is not authority to persist personal data.
Layer 4: Semantic memory
Stable facts about the user that should always be available. “Alice is the CEO of Acme. Her preferred communication style is concise. She works in Tallinn timezone.”
These are smaller in volume than episodes but more frequently retrieved. They form the agent’s “model of the user.”
Implementation: a structured profile.
{
"user_id": "alice",
"profile": {
"name": "Alice Tamm",
"role": "CEO at Acme Corp",
"location": "Tallinn, Estonia",
"timezone": "Europe/Tallinn",
"preferred_language": "English",
"communication_style": "concise, direct, no preamble",
"expertise_areas": ["product strategy", "go-to-market"],
"tools_used": ["Notion", "Slack", "Linear"]
}
}
Updates happen when the agent learns new facts. After a session, an LLM identifies new stable facts and proposes them; either auto-merged or queued for review.
Importantly: semantic facts should be confident and stable. An off-hand comment in one conversation (“I might try Python”) shouldn’t become a semantic fact (“Alice prefers Python”). The bar is higher.
A purely illustrative confidence workflow, which still requires provenance and calibration:
- Inferred once: candidate only, with the source span and no automatic behavioral effect.
- Explicitly stated: candidate for the stated scope; confirm before consequential reuse.
- Explicitly confirmed: stored with provenance, scope, review date, and user controls.
This prevents the agent from “learning” wrong facts from offhand comments.
Layer 5: Procedural memory
How the agent should behave for this user. Workflows, templates, preferences for specific actions.
Examples:
{
"user_id": "alice",
"procedural": {
"email_signature": "...",
"meeting_preferences": "always offer 3 time slots, never schedule before 9am",
"code_style": "Python, type hints required, dataclasses over dicts",
"tone_for_clients": "warm, direct, with explicit next steps",
"approval_process": "all customer-facing communications need Alice's review before sending"
}
}
These are patterns the agent follows when relevant tasks come up.
Updates may start from an explicit instruction or a repeated pattern. Repeated behavior can trigger a confirmation request, but should not silently create a durable procedure.
Storage choices
Where does memory live?
SQL database. Reliable, queryable, well-understood. Each memory type a table. Joins for retrieval. Good for structured access patterns.
Vector database. For semantic retrieval of episodes (“find memories related to this topic”). Episodes are embedded; retrieval by similarity.
Combination. SQL plus a vector index is one candidate when both structured and semantic access are required. The duplicate representation increases synchronization and deletion obligations, so benchmark it against a simpler store.
Specialized memory tools. Mem0, Letta (formerly MemGPT), Zep. These are purpose-built memory layers for agents. Worth considering if you want a higher-level abstraction.
Start with the smallest store that satisfies structured access, semantic retrieval, tenant isolation, provenance, correction, retention, deletion, backup, and recovery tests. SQL, a vector index, both, or a specialized layer may fit; compare operational and migration burden rather than assuming a team default.

Retrieval patterns
How does the agent get memory into context?
Pattern 1: Auto-load on session start
When a new session begins, automatically pull:
- The user’s semantic profile.
- The most recent N session summaries.
- Any open commitments or follow-ups.
This is the baseline context the agent has when the user shows up.
Pattern 2: Query-driven retrieval
When the user’s message hints at past topics, retrieve relevant episodes.
Example: user asks “what was the conclusion of our database discussion?” The agent searches episodes for “database” and retrieves the relevant one.
Implementation: embed the user’s message, find similar episodes, include them in context.
Pattern 3: Explicit memory tools
The agent has tools to query memory:
search_episodes(query): find specific past events.get_user_profile(): pull the semantic profile.list_open_items(): pending commitments.
The agent decides when to call these based on the conversation.
Pattern 4: Background memory enrichment
A background process periodically reviews memory and:
- Consolidates related episodes into themes.
- Updates confidence on facts.
- Decays old memories that haven’t been accessed.
This is “memory maintenance” — keeping the memory store useful over time.
Writing memory
When does memory get written?
Checkpoint or end-of-session extraction
One batch pattern, subject to durable-delivery and abrupt-termination tests. At an approved checkpoint or session end:
- An LLM analyzes the conversation.
- Extracts:
- Session summary.
- Notable events (for episodic memory).
- New facts (for semantic memory).
- Preference signals (for procedural memory).
- Updates and stores.
Batch processing can reduce in-session work, but it can also lose updates when a session ends unexpectedly and can delay correction. Measure both behavior and use a durable job where persistence is required.
Prompt for extraction:
Analyze this conversation. Output JSON with:
1. summary: 2-3 sentence summary of what happened.
2. notable_events: array of significant events worth remembering (decisions made, milestones, important context).
3. new_facts: array of stable facts learned about the user (only include if you have high confidence).
4. preference_signals: array of preferences observed (only if expressed clearly or repeated).
5. open_items: array of unresolved items the user might want to revisit.
Be conservative. Only include items with high confidence. Better to miss something than to hallucinate.
Real-time updates for high-value facts
For some facts, waiting until session end is wrong. If a user says “actually, my name is Alex, not Alice” — the correction should be applied immediately.
A pattern: have the agent detect explicit corrections or important new facts in real-time, and update memory inline.
This needs careful design — the LLM might “learn” wrong facts. Some teams require user confirmation before applying real-time updates.
User-initiated updates
The user can explicitly tell the agent things to remember:
- “Please remember that I prefer X.”
- “Forget what I said about Y.”
- “Always do Z.”
These should be first-class controls. Start the requested action immediately, show its scope and status, and explain any lawful retention or backup expiry that prevents an instant universal deletion claim. An explicit statement is strong provenance, not proof that every inferred scope is correct.
A specific tool the agent can offer:
remember(content: string, type: "fact" | "preference" | "procedure")
forget(content: string)
list_what_you_remember()
Giving users this control builds trust.
Forgetting and decay
Unbounded memory creates retrieval, cost, privacy, and accuracy risks. A documented lifecycle is essential; time-based decay is one option, not a substitute for required retention or deletion.
Time-based decay
Older memories are less likely to be retrieved. Implementation:
- Score retrieval by
relevance * recency_decay. - Old memories effectively disappear unless explicitly referenced.
Importance-based retention
Important episodes are retained longer; trivial ones decay faster.
- Tag episodes with importance at write time.
- High-value events: a purpose-specific retention period with an owner and review date; do not default to indefinite retention.
- Routine events: decay over months.
User-initiated forgetting
The user can request specific memories be deleted.
- Specific facts.
- Specific time periods.
- Specific topics.
Implementation: a deletion workflow that removes the primary record and all derived chunks, embeddings, summaries, indexes, caches, exports, and queued jobs. A tombstone may prevent re-ingestion, but merely hiding a record is not deletion. Define how backups expire and test that deleted data does not reappear after restore.
Compliance-driven deletion
Legal requirements may mandate erasure or retention. GDPR Article 17 defines a right to erasure with exceptions; legal counsel should map it and other applicable rules to the product, jurisdiction, and data role.
- User account deletion request → invoke the reviewed deletion/restriction workflow across covered stores and disclose exceptions or backup expiry.
- Per-request data deletion → locate the covered records and derivatives, then verify the result.
- Retention limits → automatically expire covered records according to the approved schedule.
These must be built in from the start. Retrofitting is painful.
Privacy considerations
Memory is sensitive. The store holds a lot about the user. Considerations:
Encryption at rest
Memory data encrypted. Standard practice.
Access controls
Who can see a user’s memory? Just them, just the system, support staff under certain conditions? Define clearly. Audit access.
PII handling
Personally identifiable information (real names, addresses, financial info) should be tagged and treated carefully. Special access controls, special deletion procedures.
User visibility
Where the product and applicable rights require it, provide a way for users to see, correct, scope, and delete memory records. A memory dashboard is one implementation; test comprehension and protect it like any other sensitive-data surface.
What the agent remembers about you:
Profile:
- Name: Alice Tamm
- Role: CEO at Acme Corp
- Communication style: concise, direct
Recent sessions:
- 2026-05-14: Drafted proposal for Acme
- 2026-05-12: Reviewed Q1 results
- ...
Preferences:
- Prefers concise responses
- Uses Notion, Slack, Linear
[Edit] [Delete specific items] [Delete all]
This transparency builds trust. Opaque memory without user visibility does not.
Sharing across contexts
If the user has multiple “modes” (work agent, personal agent), they may want memories separated. Don’t auto-share across modes unless asked.
Common failure modes
A few patterns:
Failure 1: Hallucinated memories
The agent claims to remember things that didn’t happen. “Last week we agreed on X” — but X was never discussed.
Cause: LLM “filling in” plausible-sounding memories during extraction or retrieval.
Fix: ground memory operations on real conversation data. The LLM extracts; verification is against the actual transcript. Hallucinated facts should be flagged.
Failure 2: Wrong facts learned
The agent confidently states wrong facts. “You said you prefer Python” when you actually said you were forced to use Python at work.
Cause: misinterpretation during extraction.
Fix: confidence thresholds. Only learn from explicit, repeated, or confirmed statements. User can correct.
Failure 3: Privacy leaks
Memory from one user surfaces in another’s conversation. Catastrophic.
Cause: bugs in user-scoping logic.
Fix: enforce user scoping at the storage and retrieval layers. Audit. Never trust the LLM to filter.
Failure 4: Memory bloat
After a year, memory is megabytes per user. Retrieval slows. Costs grow.
Cause: no decay or pruning.
Fix: aggressive decay. Most memory becomes inaccessible (low retrieval priority) after months. Compaction periodically.
Failure 5: Stale facts
The user changed roles 6 months ago. The agent still references the old role.
Cause: facts not updated when superseded.
Fix: detect contradictions, preserve provenance and effective dates, and ask for confirmation when the authoritative value is unclear. “Newest wins” is unsafe for delayed, quoted, or malicious input.
Failure 6: Disorienting consolidation
Background memory consolidation occasionally rewrites memories in ways that lose information.
Cause: aggressive summarization without preserving key facts.
Fix: consolidation must preserve facts explicitly. Test consolidation on real memory transcripts.
An implementation sketch: personal assistant with memory
An illustrative reference design: a personal AI assistant for individual users.
Memory layers:
- Working: current conversation.
- Session: last 7 sessions in summary form.
- Episodic: 100 most recent notable events, with semantic search.
- Semantic: user profile (name, role, preferences, tools).
- Procedural: explicit workflows the user has set up.
Storage:
- SQL (Postgres): structured profile, sessions, episodes, procedures.
- Vector DB (pgvector): episodic semantic search.
Operations:
- Session start: auto-load semantic profile + last 3 sessions + open items.
- Mid-session: episodic retrieval triggered by topic relevance.
- Session end: LLM-based extraction; user can review what was learned.
- Background: weekly consolidation (combine related episodes, decay stale ones).
User controls:
- Memory dashboard showing what’s remembered.
- Edit/delete individual items.
- “Forget the last hour” button.
- Full-account deletion workflow with verified coverage, documented exceptions, and backup-expiry behavior.
Evidence required before calling this successful:
- task completion with and without retrieved memory on a fixed evaluation set,
- precision of stored facts and retrieved memories, including contradiction handling,
- cross-tenant and cross-user isolation tests,
- correction and deletion propagation through records, embeddings, caches, exports, jobs, and backup expiry,
- token, storage, latency, and operational cost from actual traces,
- user comprehension and control testing rather than assumed confidence.
Failure modes addressed:
- Hallucinated-memory cases are included in extraction and retrieval tests.
- Confidence scores are calibrated; they do not by themselves make a fact true.
- Authorization is enforced before retrieval and again before presentation.
- Retention and consolidation jobs have audit logs, failure handling, and deletion tests.
This is a design checklist, not proof of production readiness. Production status requires implementation evidence and security/privacy review.
Specialized tools
A note on memory-as-a-service options:
Mem0. An open-source memory layer. Evaluate its current documentation and code against your persistence, isolation, and deletion requirements.
Letta (MemGPT). A tool-oriented memory design. Review its current documentation and operational boundaries.
Zep. A hosted memory and context service. Validate its current documentation, data boundary, and deletion contract.
Cognee. A knowledge-graph-oriented option. Validate maturity and fit from its current documentation and repository before adopting it.
These tools may reduce implementation work and add vendor, security, migration, and data-lifecycle dependencies. Compare them with an internal design using the same acceptance tests.
Build the minimum justified memory
Long-term memory is what makes agents continuous across sessions rather than amnesiac. It’s also one of the harder things to get right.
The architecture is layered:
- Working memory (in-context).
- Session memory (recent sessions).
- Episodic memory (specific events).
- Semantic memory (stable facts).
- Procedural memory (preferences and workflows).
Each layer has its own storage, retrieval, and decay logic. Each contributes to making the agent useful across time.
The patterns that matter:
- Conservative extraction (don’t hallucinate facts).
- Confidence-based learning (don’t learn from offhand comments).
- Active forgetting (decay and pruning).
- User control (transparency and edit ability).
- Privacy enforcement (at every layer).
When it passes the evaluation and lifecycle tests, memory can reduce repeated briefing and make confirmed preferences available across sessions.
For agents that require cross-session continuity, persistence is an explicit product choice. Build the minimum justified memory, with provenance, authorization, user control, and a tested end-of-life path.



