Memory is becoming a standard component in agent architectures. It's also an attack surface. Your agent reading from it doesn't know the difference.

Scenario one: an agent in a multi-tenant system. The agent pulls context from memory: previous decisions, user preferences, system state. Everything looks correct. The agent acts with full confidence. Some of those memories belong to a different tenant. Some were written from a data source that had already drifted. The agent never flags uncertainty. It just acts.

A single drop of ink spreading through clear water, edges dissolving at the boundary between what was certain and what has already changed

Scenario two: shared memory across automated tools. A code review system running across thousands of repos. Every review produces feedback. The feedback accumulates in shared memory, shaping future reviews across repos, languages, and teams. An early bad signal doesn't decay. It compounds. By the time anyone notices, the system has been approving buggy patterns, wasting cycles chasing phantom issues, or worse: accepting infiltrated code it was trained to trust.

Both scenarios share the same root failure: memory as an unguarded input surface with no isolation boundary around writes.

The agent acts on its map, not the territory. Poison the memory, corrupt the map.


Two vectors get you to the same place: accidental and adversarial.

The first is accidental. Bad source data gets ingested into memory: stale documentation, drifted configs, incorrect prior outputs, outdated state. The agent didn't ask for bad data. The pipeline gave it bad data and the memory held it.

The second is adversarial. A user or process deliberately writes false context. This takes two forms. Knowledge poisoning injects false facts that cause the agent to give bad advice in future sessions. Instruction injection is more dangerous: behavior-changing instructions written into memory that activate silently in future conversations. The adversary doesn't need access to your system prompt. They need write access to your memory store.

Both forms can arrive directly, through a user crafting inputs that get stored, or indirectly, through a compromised upstream tool writing to shared memory.

Both vectors produce the same symptom: confident action on wrong context. That shared symptom is why they require a unified architectural response, not two separate mitigations.


The response looks different depending on which architecture you're building. First: closed multi-tenant systems.

Closed multi-tenant systems have a known set of failure modes.

Shared embedding stores without tenant-scoped namespacing let one tenant's context bleed into another's. Memory that outlives session boundaries means user A's decisions are still in the store when user B starts. Writes without provenance give you no record of what wrote a memory or when. No TTLs mean stale state persists indefinitely.

Three defenses address this.

Isolation. Tenant-scoped namespaces on all reads and writes, enforced at the storage layer, not the application layer. Scope by tenant, session, and agent identity. Namespacing enforced at the application layer can be bypassed by bugs in context assembly. Enforce at the layer that doesn't trust its callers.

Provenance and TTL. Every memory entry records its source, timestamp, and the input that generated it. Memory expires by default; explicit renewal required. TTLs that expire mid-task can leave agents with incomplete context, so design session locks or grace periods for long-running operations. Rollback reverts the memory store; it doesn't undo downstream side effects already emitted.

Read audit trails. Log what the agent read before it acted. Post-incident investigation requires this. You can't trace contamination you didn't log.

Closed systems can approach elimination of poisoning with strict architecture. The hard case is semantic poisoning: wrong memories that don't contradict existing state and can't be detected without ground truth. That's what makes shared feedback loops a harder problem.


Here's what that harder problem looks like at scale.

Take scenario two. The code review system needs to learn what acceptable code looks like per language, per framework, per project. Every CI run produces signal. Thousands of runs per day. The flywheel turns.

But the flywheel has no opinion about signal quality. One team's bad pattern, accepted as feedback enough times, becomes a standard. Hierarchical shared memory means contamination at a parent node propagates to every child. The blast radius of a single poisoned node scales with your CI volume. At a few hundred runs per day, you notice the drift. At thousands, it's already policy before you see it.

Now take scenario one at scale. You're running a shared agentic orchestration system, isolated per tenant. Each tenant's decisions, process outcomes, and agent behavior are scoped to their context. But you want the system to improve. You want tenant signal to feed back into central processes, surface better patterns, optimize decisions downstream.

To do that, you have to let signal cross the isolation boundary.

That crossing is exactly where the attack surface opens. You can scope reads. You can isolate writes. But when you want the flywheel to turn, you have to accept signal from tenants. An adversary doesn't need to break isolation. They just need to be a tenant.


In a feedback loop, bad signal doesn't get corrected. It gets reinforced. By the time poisoning is detectable, the contaminated memories may have been cited enough times to feel authoritative.

In feedback loops, the defenses look different because the threat model is different.

Write authorization and signing. Each tool, agent, or tenant has a signing identity. Memory writes are attributed and only accepted from authorized sources. A compromised or adversarial source can be revoked without rolling back the entire store. The key property is revocability: you need to be able to excise a source without losing everything else.

Feedback gates. Outputs feeding back into durable memory require a confidence threshold or human review gate before acceptance. Without this, the loop amplifies noise as readily as signal. Gates add friction. That friction is the feature.

Versioned snapshots and rollback. Checkpoint memory state at intervals. When bad feedback is detected, rollback is possible. Immutable history means contamination can be traced to its source and excised, not just overwritten.

Contradiction detection helps. When a new write conflicts with existing memory, surfacing the conflict is better than silently overwriting. But it isn't sufficient alone. Semantic poisoning, wrong memories that don't obviously contradict existing state, requires gates.

You can't make an open feedback loop fully safe without gates. Gates slow the loop. What confidence threshold justifies crossing the isolation boundary is a product decision, not a technical one.


When memory gets treated as a feature rather than infrastructure, the trust boundary never gets built. Spin the flywheel without it and you don't get a smarter system. You get a faster one. Faster at acting on a poisoned map at scale.

The agent doesn't know its map is wrong. That's the point. Build the infrastructure that keeps the map accurate before you need it to be.