Most agent frameworks ship a single object called memory. You attach a vector store, the framework embeds the conversation, and retrieval happens on every turn. That design collapses three unrelated problems with three unrelated lifecycles into one similarity search, and it is the reason agents that demo well fall apart on day-three workflows.
Here is the direct answer. Agent memory for long-running enterprise workflows is not one store. It is four separable layers: working context (tokens in the model window right now), durable run state (the checkpointed record of what the agent decided and did), episodic memory (completed runs with outcome labels), and shared organizational knowledge (policy, product, and account data that no single run owns). Each layer needs its own write path, its own retention rule, and its own owner. A vector database serves one of the four.
The failure mode is specific. Retrieval similarity has no concept of recency authority. A superseded decision recorded at step 4 can outrank the corrected decision recorded at step 40, because the superseded text happens to be a closer semantic match to the current question. Once an agent runs for minutes or days, mutates real systems through tools, and survives process restarts, you are operating a long-lived stateful process, which pulls in failure-handling and recovery concerns that the AWS Well-Architected reliability guidance addresses directly and that prompt engineering never did [2]. This article covers the architecture for state that survives restarts, respects retention policy, and does not quietly accumulate stale context.
A Vector Store Is Not Memory
Retrieval answers one question: what text in my corpus resembles this query? That is genuinely useful for shared organizational knowledge. It is the wrong mechanism for the other three layers.
Durable run state needs exact reads, not approximate ones. When an agent resumes after an eviction, it must know precisely which step it completed, which tool calls already produced side effects, and which approvals are still pending. Approximate retrieval over a transcript cannot answer that with the certainty a write-path decision requires.
Episodic memory needs outcome labels and time bounds. "Find three similar past runs" is only useful if you also know whether those runs succeeded, what the human reviewer corrected, and whether the policy that governed them still applies. Similarity alone will happily hand the model three examples of the wrong behaviour.
Working context needs aggressive discipline in the opposite direction. The window should be reconstructed from durable state at each step, not accumulated by appending everything that happened. Accumulation is how token budgets blow out and how contradictions enter the same prompt.
Four Memory Layers, Four Different Lifecycles
Separating the layers changes what each one is allowed to do. Working context is ephemeral by design and may be discarded at any time without loss, because it is derived. Durable run state is authoritative and must be written before the agent acts on it. Episodic memory is append-mostly and read at planning time. Shared knowledge is owned by data stewards outside the agent, and no run should mutate it.
Approval and audit records deserve their own row, because they have a different consumer. Engineers read run state to debug. Auditors and risk owners read the approval trail to reconstruct who authorized an irreversible action and when. Keeping the approval decision only in a chat thread means it does not exist as evidence.
| Layer | What it stores | Write path | Retention class | Failure if wrong |
|---|---|---|---|---|
| Working context | Tokens assembled for the current step | Derived, never authoritative | Discard at step end | Context bloat, contradictory facts in one prompt |
| Durable run state | Step index, committed facts, pending approvals, tool results by idempotency key | Checkpoint write after every step | Retain to run close plus audit window | Restart repeats a write side effect |
| Episodic memory | Completed run summaries with outcome labels | Batch write at run close | Tiered expiry by use case | Model grounds on stale or wrong-outcome examples |
| Shared knowledge | Policy, product, entity, account data | Owned by data stewards, agent reads only | Follows source system retention | Agent cites retired policy as current |
| Approval and audit | Decision, approver identity, timestamp, scope granted | Immutable append at gate crossing | Longest class, legal or regulatory driver | No defensible record of who authorized the action |
The most common architectural mistake is letting one store serve rows two and three. Run state then inherits episodic retention, episodic memory inherits run-state write frequency, and deletion becomes impossible at the record level because everything is entangled in one index.
Checkpoint the Decision, Not the Transcript
The rule is simple: persist structured state after every step so a failed or evicted run resumes from the last good state instead of restarting and repeating side effects [2]. What you checkpoint matters as much as when.
Transcript-append checkpointing fails three ways. It grows without bound, so resumption cost rises with run length. It mixes reasoning noise with committed facts, so nothing downstream can tell a hypothesis from a decision. And it cannot be diffed or reviewed by a human, which means no one can answer "what did the agent believe before this step?" during an incident.
Checkpoint the decision instead:
{
"run_id": "run_8f21c4",
"step_index": 17,
"stage": "awaiting_credit_approval",
"committed_facts": [
{ "key": "account_id", "value": "ACC-44190", "source": "crm.get_account", "valid_from": "2026-03-04T11:02Z" },
{ "key": "outstanding_balance", "value": "18400.00", "source": "erp.get_ar", "valid_from": "2026-03-04T11:06Z", "supersedes": "fact_9c1" }
],
"open_questions": ["tax_jurisdiction_unconfirmed"],
"pending_approvals": [
{ "action": "issue_credit_memo", "amount": "18400.00", "requested_at": "2026-03-04T11:09Z", "state": "pending", "gate": "finance_controller" }
],
"tool_results": {
"erp.post_credit_memo:idem_7b22a9": { "status": "not_attempted" }
},
"budget": { "steps_used": 17, "step_ceiling": 60, "cost_used_usd": 1.94, "cost_ceiling_usd": 12.00 }
}Three properties make this checkpoint useful. Facts carry provenance and a supersede pointer, so precedence is decidable without asking the model. Pending approvals are state, not conversation. And tool results are keyed by idempotency key, which pairs the memory design with the write-side control that actually prevents duplicated side effects.
That pairing is the part teams skip. Classify every tool as read or write, then require an idempotency key plus a server-side dedupe window on every write tool, so a retried step cannot duplicate the effect [2]. Memory tells the agent "I already called this"; the dedupe window enforces it even when memory is stale. You need both, because at-least-once delivery means the retry may arrive before your checkpoint write lands.
Where Stale Context Corrupts Decisions
Stale memory rarely produces an obvious error. It produces a confidently wrong decision that looks reasonable in the trace.
Three corruption patterns account for most of what I see in production reviews. Superseded facts retrieved as current, where a quoted price from early in the run outranks the renegotiated price. Stale entity state, where a closed account is still described as active because the summary written at step 6 said so. Contradictory duplicates, where repeated summarization produces two versions of the same fact with no tiebreaker.
The fix is a precedence rule enforced at read time and a reconciliation rule enforced at write time. Every memory record carries a validity window and an optional supersede pointer. Retrieval filters on validity before it ranks on similarity:
def recall(query, run_id, as_of):
candidates = store.search(query, run_id=run_id, top_k=40)
live = [
r for r in candidates
if r.valid_from <= as_of
and (r.valid_to is None or r.valid_to > as_of)
and r.superseded_by is None
]
return rank_by_similarity(live, query)[:8]Prefer write-time discipline over read-time cleverness. Reconciling a contradiction when the fact is written is cheaper, deterministic, and auditable. Asking the model to referee two contradictory records at read time turns a data-integrity problem into a coin flip you cannot test.
Then instrument what entered the prompt. Log the record IDs assembled into each step's context alongside the run ID and step index. Without that log, a bad decision is untraceable: you can see the output but not the input that caused it. With it, you can walk from the wrong credit memo back to the specific superseded record that should have been filtered.
Retention Policy Is a Memory Architecture Constraint
Retention cannot be added after launch, because deletion has to work at the record level across every derived store. That is an architectural property, not a policy document.
The derived-copy problem is where most retention programs quietly fail. Deleting a source document does not delete its embedding in the vector index, the cached summary that quoted it, the episodic run record that grounded on it, or the trace that logged its text into a prompt. Unless each derived artifact carries a pointer back to its source, a deletion request becomes a manual archaeology exercise.
Treat the governance side as mapping, not new bureaucracy. Tier agent use cases by consequence so an internal summarizer and an agent that issues credit memos face proportionate review depth, using the function-based structure of the NIST AI Risk Management Framework as the organizing spine [3]. Where the organization wants auditable repeatability, ISO/IEC 42001 sets out requirements for an AI management system that can be operated like other management systems [4]. Attach the evidence to artifacts engineers already produce: pull requests, evaluation reports, deployment records [3].
Memory-specific re-review triggers matter more than annual cycles. Re-review on expanded autonomy scope, on new tool permissions, and on any change to what the agent is allowed to remember across runs [3]. That last trigger is the one teams forget, and it is the one that changes the data-protection profile most.
| Record type | Example | Retention driver | Deletion path | Review trigger |
|---|---|---|---|---|
| Working context | Assembled prompt for step 17 | None, transient | Dropped at step end; scrub from trace sink | Any change to what is logged |
| Run state | Checkpoint with committed facts | Operational plus audit window | TTL on run close plus retention period | New write-tool permission |
| Episodic memory | Closed-run summary with outcome label | Use-case tier and data category | Record-level delete plus embedding purge | Cross-run memory scope change |
| Derived embeddings | Vector for a policy chunk | Inherits source document | Cascade delete keyed by source ID | Source system retention change |
| Approval trail | Controller approval of credit memo | Legal or regulatory | Restricted; legal-hold aware | Autonomy scope expansion |
Patterns That Hold Up in Production
Three patterns survive contact with real workflows.
State machine with model-chosen transitions. The stages and allowed transitions are code. The agent chooses which allowed transition to take. Current stage lives in durable state rather than being inferred from the transcript, which means resumption reads a field instead of re-deriving intent from prose.
Blackboard for multi-agent handoff. Agents write typed, schema-validated artifacts into a shared versioned workspace instead of passing prose context to each other. Prose handoff loses constraints at every hop; a typed artifact either validates or it does not. This is the single highest-value change for teams whose multi-agent systems fail in ways nobody can reproduce.
Memory as a tool. The agent calls explicit memory.read and memory.write tools rather than having state injected invisibly by the framework. Every state change then appears in the trace with its arguments, which makes it gateable, deniable, and reviewable [2]. It also means your memory writes get the same idempotency treatment as any other write tool.
Human approval belongs in persisted state across all three. The decision, the approver identity, the scope granted, and the timestamp go in the run record, not only in a Slack thread [2].
| Pattern | State store | Handoff mechanism | Audit strength | Best for |
|---|---|---|---|---|
| State machine | Checkpointed stage plus facts | Stage transition events | High: stage history is explicit | Regulated, ordered workflows with approval gates |
| Blackboard | Versioned artifact workspace | Typed schema-validated artifacts | High: artifact versions are diffable | Multi-agent research, underwriting, case assembly |
| Memory as tool | Whatever backs the tool | Explicit read and write calls | Highest: every write is a traced call | Any workflow needing per-write authorization |
| Framework default memory | Opaque conversation store | Implicit context injection | Low: writes are invisible | Prototypes and single-session assistants |
Most of our agentic AI engagements start with the state machine because it produces the fewest surprises, then add a blackboard when a second agent enters the workflow.
Testing Memory: Replay Beats Unit Tests
Memory bugs do not appear in single-turn tests. They emerge across step boundaries, after restarts, and after compaction. That makes the full task episode the unit of evaluation: the prompt, every tool call and argument, every tool response, and the end state of the environment [1].
Build a replay harness. Record production runs as trajectories, then re-execute them against candidate prompts, tool schemas, and model versions before rollout [2]. This is the only practical way to catch a memory-driven behavioural regression, because the regression is a change in path, not a change in final answer. Two versions can produce the same output while one of them called the write tool twice.
Prescribe memory-specific cases:
- Kill and resume mid-run. Terminate the process between a tool call and its checkpoint write, then assert the resumed run does not repeat the side effect.
- Inject a duplicate tool response. Deliver the same tool result twice with the same idempotency key and assert state converges rather than double-counting.
- Supersede a fact mid-run. Change an upstream value after the agent has cached it and assert the newer record wins on the next read.
- Expire a record the agent already used. Assert the agent either refreshes it or stops, rather than silently continuing on an expired fact.
- Feed instruction-like text through memory. Retrieved content and tool output containing directives must be treated as data, not as a new instruction [1].
Wire the suite into CI as a release gate that runs on every prompt, tool-schema, and model-version change, not only at milestones [1]. Our agent harness design checklist covers how the gate fits alongside budget ceilings and trace instrumentation.
FAQ and Where to Start This Week
Is a vector database enough for agent memory? No. It serves shared organizational knowledge well. It does not provide the exact reads, precedence rules, or idempotency records that durable run state requires.
Do I need a durable execution engine? You need durable checkpoints and resumability [2]. An engine is one way to get them. A well-designed state table with per-step writes and idempotency keys is another. Choose based on how many concurrent long runs you expect and who operates the platform.
How long should episodic memory live? Tier it by use-case consequence and data category [3]. Runs that grounded on regulated personal data should expire faster than runs about internal tooling, and every episodic record needs a deletion path that also purges its embeddings.
Who owns shared organizational knowledge? The source-system data steward, not the agent team. Agents read it. If an agent needs to write organizational facts, route that write through the owning system so its retention and review rules apply.
What belongs in the audit trail? Run ID, step index, tool name, arguments, latency, outcome, and every approval decision with approver and timestamp [2].
Your next working session
Open your agent's memory module and label every write as working context, run state, episodic, or shared. Assign each one a TTL and a named owner. Any write you cannot classify is a bug, and any write without an owner will become a retention finding later.
Then start tracking one metric this week: resume success rate, the share of interrupted runs that continue from checkpoint without repeating a write-side effect. It is the cheapest single signal that your memory architecture is real rather than aspirational.
Record the choice as a short architecture decision record stating the options, the trade-offs accepted, and the conditions that would trigger revisiting it, reviewed against the Well-Architected pillars so reliability and cost questions get asked consistently [5].
The framework that gave you one memory object was not wrong to make it easy. It was wrong to make it singular. Four layers, four lifecycles, four owners.
References
[1]Liu et al., "AgentBench: Evaluating LLMs as Agents," arXiv preprint. https://arxiv.org/abs/2308.03688
[2]AWS, "AWS Well-Architected Reliability Pillar," AWS Documentation. https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html
[3]NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1." https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
[4]ISO/IEC, "ISO/IEC 42001 Artificial intelligence management system standard." https://www.iso.org/standard/81230.html
[5]AWS, "AWS Well-Architected Framework," AWS Documentation. https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html