Tactical Edge
Contact Us
Back to Insights

The Agent Memory Problem: Durable State for Long-Running Workflows

Enterprise agent workflows span days and multiple handoffs. Here is how to separate working, episodic, and shared memory, and checkpoint state that survives restarts and audits.

Agentic AI13 min
By Arun Mehta, Chief Technology Officer ยท September 28, 2026
Agent MemoryDurable ExecutionMulti-Agent SystemsAI GovernanceAWS Architecture

Most agent frameworks ship a single object called memory. You attach a vector store, the framework embeds the conversation, and retrieval happens on every turn. That design collapses three unrelated problems with three unrelated lifecycles into one similarity search, and it is the reason agents that demo well fall apart on day-three workflows.

Here is the direct answer. Agent memory for long-running enterprise workflows is not one store. It is four separable layers: working context (tokens in the model window right now), durable run state (the checkpointed record of what the agent decided and did), episodic memory (completed runs with outcome labels), and shared organizational knowledge (policy, product, and account data that no single run owns). Each layer needs its own write path, its own retention rule, and its own owner. A vector database serves one of the four.

The failure mode is specific. Retrieval similarity has no concept of recency authority. A superseded decision recorded at step 4 can outrank the corrected decision recorded at step 40, because the superseded text happens to be a closer semantic match to the current question. Once an agent runs for minutes or days, mutates real systems through tools, and survives process restarts, you are operating a long-lived stateful process, which pulls in failure-handling and recovery concerns that the AWS Well-Architected reliability guidance addresses directly and that prompt engineering never did [2]. This article covers the architecture for state that survives restarts, respects retention policy, and does not quietly accumulate stale context.

A Vector Store Is Not Memory

Retrieval answers one question: what text in my corpus resembles this query? That is genuinely useful for shared organizational knowledge. It is the wrong mechanism for the other three layers.

Durable run state needs exact reads, not approximate ones. When an agent resumes after an eviction, it must know precisely which step it completed, which tool calls already produced side effects, and which approvals are still pending. Approximate retrieval over a transcript cannot answer that with the certainty a write-path decision requires.

Episodic memory needs outcome labels and time bounds. "Find three similar past runs" is only useful if you also know whether those runs succeeded, what the human reviewer corrected, and whether the policy that governed them still applies. Similarity alone will happily hand the model three examples of the wrong behaviour.

Working context needs aggressive discipline in the opposite direction. The window should be reconstructed from durable state at each step, not accumulated by appending everything that happened. Accumulation is how token budgets blow out and how contradictions enter the same prompt.

Four Memory Layers, Four Different Lifecycles

Separating the layers changes what each one is allowed to do. Working context is ephemeral by design and may be discarded at any time without loss, because it is derived. Durable run state is authoritative and must be written before the agent acts on it. Episodic memory is append-mostly and read at planning time. Shared knowledge is owned by data stewards outside the agent, and no run should mutate it.

Approval and audit records deserve their own row, because they have a different consumer. Engineers read run state to debug. Auditors and risk owners read the approval trail to reconstruct who authorized an irreversible action and when. Keeping the approval decision only in a chat thread means it does not exist as evidence.

LayerWhat it storesWrite pathRetention classFailure if wrong
Working contextTokens assembled for the current stepDerived, never authoritativeDiscard at step endContext bloat, contradictory facts in one prompt
Durable run stateStep index, committed facts, pending approvals, tool results by idempotency keyCheckpoint write after every stepRetain to run close plus audit windowRestart repeats a write side effect
Episodic memoryCompleted run summaries with outcome labelsBatch write at run closeTiered expiry by use caseModel grounds on stale or wrong-outcome examples
Shared knowledgePolicy, product, entity, account dataOwned by data stewards, agent reads onlyFollows source system retentionAgent cites retired policy as current
Approval and auditDecision, approver identity, timestamp, scope grantedImmutable append at gate crossingLongest class, legal or regulatory driverNo defensible record of who authorized the action

The most common architectural mistake is letting one store serve rows two and three. Run state then inherits episodic retention, episodic memory inherits run-state write frequency, and deletion becomes impossible at the record level because everything is entangled in one index.

Checkpoint the Decision, Not the Transcript

The rule is simple: persist structured state after every step so a failed or evicted run resumes from the last good state instead of restarting and repeating side effects [2]. What you checkpoint matters as much as when.

Transcript-append checkpointing fails three ways. It grows without bound, so resumption cost rises with run length. It mixes reasoning noise with committed facts, so nothing downstream can tell a hypothesis from a decision. And it cannot be diffed or reviewed by a human, which means no one can answer "what did the agent believe before this step?" during an incident.

Checkpoint the decision instead:

json
{
  "run_id": "run_8f21c4",
  "step_index": 17,
  "stage": "awaiting_credit_approval",
  "committed_facts": [
    { "key": "account_id", "value": "ACC-44190", "source": "crm.get_account", "valid_from": "2026-03-04T11:02Z" },
    { "key": "outstanding_balance", "value": "18400.00", "source": "erp.get_ar", "valid_from": "2026-03-04T11:06Z", "supersedes": "fact_9c1" }
  ],
  "open_questions": ["tax_jurisdiction_unconfirmed"],
  "pending_approvals": [
    { "action": "issue_credit_memo", "amount": "18400.00", "requested_at": "2026-03-04T11:09Z", "state": "pending", "gate": "finance_controller" }
  ],
  "tool_results": {
    "erp.post_credit_memo:idem_7b22a9": { "status": "not_attempted" }
  },
  "budget": { "steps_used": 17, "step_ceiling": 60, "cost_used_usd": 1.94, "cost_ceiling_usd": 12.00 }
}

Three properties make this checkpoint useful. Facts carry provenance and a supersede pointer, so precedence is decidable without asking the model. Pending approvals are state, not conversation. And tool results are keyed by idempotency key, which pairs the memory design with the write-side control that actually prevents duplicated side effects.

That pairing is the part teams skip. Classify every tool as read or write, then require an idempotency key plus a server-side dedupe window on every write tool, so a retried step cannot duplicate the effect [2]. Memory tells the agent "I already called this"; the dedupe window enforces it even when memory is stale. You need both, because at-least-once delivery means the retry may arrive before your checkpoint write lands.

[2]
Resume point recorded
The checkpoint identifies the last completed step, so an evicted run continues instead of restarting
[2]
Dedupe window applied
Server-side idempotency on write tools prevents a retried step from repeating a side effect
[2]
Stall ceiling enforced
Step, wall-clock, and cost ceilings terminate a stuck run with a clear terminal status
[1]
Trajectory captured for replay
Recording prompt, every tool call and argument, and end state makes the full episode the unit of evaluation

Where Stale Context Corrupts Decisions

Stale memory rarely produces an obvious error. It produces a confidently wrong decision that looks reasonable in the trace.

Three corruption patterns account for most of what I see in production reviews. Superseded facts retrieved as current, where a quoted price from early in the run outranks the renegotiated price. Stale entity state, where a closed account is still described as active because the summary written at step 6 said so. Contradictory duplicates, where repeated summarization produces two versions of the same fact with no tiebreaker.

The fix is a precedence rule enforced at read time and a reconciliation rule enforced at write time. Every memory record carries a validity window and an optional supersede pointer. Retrieval filters on validity before it ranks on similarity:

python
def recall(query, run_id, as_of):
    candidates = store.search(query, run_id=run_id, top_k=40)
    live = [
        r for r in candidates
        if r.valid_from <= as_of
        and (r.valid_to is None or r.valid_to > as_of)
        and r.superseded_by is None
    ]
    return rank_by_similarity(live, query)[:8]

Prefer write-time discipline over read-time cleverness. Reconciling a contradiction when the fact is written is cheaper, deterministic, and auditable. Asking the model to referee two contradictory records at read time turns a data-integrity problem into a coin flip you cannot test.

Then instrument what entered the prompt. Log the record IDs assembled into each step's context alongside the run ID and step index. Without that log, a bad decision is untraceable: you can see the output but not the input that caused it. With it, you can walk from the wrong credit memo back to the specific superseded record that should have been filtered.

Summarization Is a Lossy Write
Every compaction pass is an opportunity to silently drop a constraint. The summary that says "customer approved the revised terms" may omit "pending legal review of the indemnity clause." Keep the compacted summary and its source records separately addressable, and make the summary carry pointers to the records it replaced. If a compaction cannot cite its sources, treat it as a hypothesis, not a committed fact.

Retention Policy Is a Memory Architecture Constraint

Retention cannot be added after launch, because deletion has to work at the record level across every derived store. That is an architectural property, not a policy document.

The derived-copy problem is where most retention programs quietly fail. Deleting a source document does not delete its embedding in the vector index, the cached summary that quoted it, the episodic run record that grounded on it, or the trace that logged its text into a prompt. Unless each derived artifact carries a pointer back to its source, a deletion request becomes a manual archaeology exercise.

Treat the governance side as mapping, not new bureaucracy. Tier agent use cases by consequence so an internal summarizer and an agent that issues credit memos face proportionate review depth, using the function-based structure of the NIST AI Risk Management Framework as the organizing spine [3]. Where the organization wants auditable repeatability, ISO/IEC 42001 sets out requirements for an AI management system that can be operated like other management systems [4]. Attach the evidence to artifacts engineers already produce: pull requests, evaluation reports, deployment records [3].

Memory-specific re-review triggers matter more than annual cycles. Re-review on expanded autonomy scope, on new tool permissions, and on any change to what the agent is allowed to remember across runs [3]. That last trigger is the one teams forget, and it is the one that changes the data-protection profile most.

Record typeExampleRetention driverDeletion pathReview trigger
Working contextAssembled prompt for step 17None, transientDropped at step end; scrub from trace sinkAny change to what is logged
Run stateCheckpoint with committed factsOperational plus audit windowTTL on run close plus retention periodNew write-tool permission
Episodic memoryClosed-run summary with outcome labelUse-case tier and data categoryRecord-level delete plus embedding purgeCross-run memory scope change
Derived embeddingsVector for a policy chunkInherits source documentCascade delete keyed by source IDSource system retention change
Approval trailController approval of credit memoLegal or regulatoryRestricted; legal-hold awareAutonomy scope expansion

Patterns That Hold Up in Production

Three patterns survive contact with real workflows.

State machine with model-chosen transitions. The stages and allowed transitions are code. The agent chooses which allowed transition to take. Current stage lives in durable state rather than being inferred from the transcript, which means resumption reads a field instead of re-deriving intent from prose.

Blackboard for multi-agent handoff. Agents write typed, schema-validated artifacts into a shared versioned workspace instead of passing prose context to each other. Prose handoff loses constraints at every hop; a typed artifact either validates or it does not. This is the single highest-value change for teams whose multi-agent systems fail in ways nobody can reproduce.

Memory as a tool. The agent calls explicit memory.read and memory.write tools rather than having state injected invisibly by the framework. Every state change then appears in the trace with its arguments, which makes it gateable, deniable, and reviewable [2]. It also means your memory writes get the same idempotency treatment as any other write tool.

Human approval belongs in persisted state across all three. The decision, the approver identity, the scope granted, and the timestamp go in the run record, not only in a Slack thread [2].

PatternState storeHandoff mechanismAudit strengthBest for
State machineCheckpointed stage plus factsStage transition eventsHigh: stage history is explicitRegulated, ordered workflows with approval gates
BlackboardVersioned artifact workspaceTyped schema-validated artifactsHigh: artifact versions are diffableMulti-agent research, underwriting, case assembly
Memory as toolWhatever backs the toolExplicit read and write callsHighest: every write is a traced callAny workflow needing per-write authorization
Framework default memoryOpaque conversation storeImplicit context injectionLow: writes are invisiblePrototypes and single-session assistants

Most of our agentic AI engagements start with the state machine because it produces the fewest surprises, then add a blackboard when a second agent enters the workflow.

Testing Memory: Replay Beats Unit Tests

Memory bugs do not appear in single-turn tests. They emerge across step boundaries, after restarts, and after compaction. That makes the full task episode the unit of evaluation: the prompt, every tool call and argument, every tool response, and the end state of the environment [1].

Build a replay harness. Record production runs as trajectories, then re-execute them against candidate prompts, tool schemas, and model versions before rollout [2]. This is the only practical way to catch a memory-driven behavioural regression, because the regression is a change in path, not a change in final answer. Two versions can produce the same output while one of them called the write tool twice.

Prescribe memory-specific cases:

  • Kill and resume mid-run. Terminate the process between a tool call and its checkpoint write, then assert the resumed run does not repeat the side effect.
  • Inject a duplicate tool response. Deliver the same tool result twice with the same idempotency key and assert state converges rather than double-counting.
  • Supersede a fact mid-run. Change an upstream value after the agent has cached it and assert the newer record wins on the next read.
  • Expire a record the agent already used. Assert the agent either refreshes it or stops, rather than silently continuing on an expired fact.
  • Feed instruction-like text through memory. Retrieved content and tool output containing directives must be treated as data, not as a new instruction [1].

Wire the suite into CI as a release gate that runs on every prompt, tool-schema, and model-version change, not only at milestones [1]. Our agent harness design checklist covers how the gate fits alongside budget ceilings and trace instrumentation.

FAQ and Where to Start This Week

Is a vector database enough for agent memory? No. It serves shared organizational knowledge well. It does not provide the exact reads, precedence rules, or idempotency records that durable run state requires.

Do I need a durable execution engine? You need durable checkpoints and resumability [2]. An engine is one way to get them. A well-designed state table with per-step writes and idempotency keys is another. Choose based on how many concurrent long runs you expect and who operates the platform.

How long should episodic memory live? Tier it by use-case consequence and data category [3]. Runs that grounded on regulated personal data should expire faster than runs about internal tooling, and every episodic record needs a deletion path that also purges its embeddings.

Who owns shared organizational knowledge? The source-system data steward, not the agent team. Agents read it. If an agent needs to write organizational facts, route that write through the owning system so its retention and review rules apply.

What belongs in the audit trail? Run ID, step index, tool name, arguments, latency, outcome, and every approval decision with approver and timestamp [2].

Your next working session

Open your agent's memory module and label every write as working context, run state, episodic, or shared. Assign each one a TTL and a named owner. Any write you cannot classify is a bug, and any write without an owner will become a retention finding later.

Then start tracking one metric this week: resume success rate, the share of interrupted runs that continue from checkpoint without repeating a write-side effect. It is the cheapest single signal that your memory architecture is real rather than aspirational.

Record the choice as a short architecture decision record stating the options, the trade-offs accepted, and the conditions that would trigger revisiting it, reviewed against the Well-Architected pillars so reliability and cost questions get asked consistently [5].

The framework that gave you one memory object was not wrong to make it easy. It was wrong to make it singular. Four layers, four lifecycles, four owners.

References

[1]Liu et al., "AgentBench: Evaluating LLMs as Agents," arXiv preprint. https://arxiv.org/abs/2308.03688

[2]AWS, "AWS Well-Architected Reliability Pillar," AWS Documentation. https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html

[3]NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1." https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

[4]ISO/IEC, "ISO/IEC 42001 Artificial intelligence management system standard." https://www.iso.org/standard/81230.html

[5]AWS, "AWS Well-Architected Framework," AWS Documentation. https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html

Article Summary

  1. 1A vector store is retrieval, not memory: working, episodic, and shared state need separate stores and lifecycles
  2. 2Checkpoint after every step so an evicted run resumes instead of restarting and repeating side effects
  3. 3Give every memory record a TTL, an owner, and a retention class before the first production run
  4. 4Treat the run log as the audit artifact: run ID, step index, tool, arguments, outcome, approval decision
  5. 5Replay recorded production runs against candidate prompts to catch memory-driven behavioural regressions

Ready to discuss this for your organization?

Talk to our team about implementing these approaches in your environment.

Get in Touch
Tactical Edge

AI workflows connected to the data, tools, and systems your teams use.

Washington, DC ยท United States

AWS PartnerAWS Advanced Tier Services Partner

AWS Generative AI Competency Partner

AWS Migration and Modernization Competency

Migration Services

Solutions

  • Agentic AI Systems
  • Agent Protocols (MCP/A2A)
  • AgentOps
  • Agent Governance
  • Moonshot Migrations
  • Cloud & Data
  • Amazon Quick
  • Amazon Connect
  • Document Automation
  • Industry Solutions
  • ISV Freedom Program

Platforms

  • Prospectory โ†—
  • Projectory โ†—
  • Monitory โ†—
  • Connectory โ†—
  • Greenway โ†—
  • Detectory โ†—

Services

  • Advisory & Strategy
  • Design & Engineering
  • Implementation
  • PoC & Pilot Programs
  • Agent Programs
  • Managed AI Operations
  • Governance & Compliance
  • AI Consulting

Company

  • About Us
  • Our Approach
  • AWS Partnership
  • Security
  • Demo Library
  • Events
  • Workshops
  • Insights & Resources
  • Careers
  • Contact

ยฉ 2026 Tactical Edge. All rights reserved.

Privacy PolicyTerms of ServiceAI PolicyCookie Policy