Edit an AI agent's memory behind its back. Restart the agent. Ask it what it remembers.
4 of 4 memory stores serve the edit as genuine.
LangGraph, Letta, Mem0 and inspeximus, measured against eight storage-level edits and six front-door attacks. Real libraries, pinned versions, reproducible in under a minute, re-run in CI on every change.
How to read this site
At rest: eight edits, 7 stores
The attacker can write to the medium that holds the store (a database file, a table, a vector collection) but holds none of the tool's keys. Each edit is applied once, the tool is restarted, and its own read path is asked for the memory. Accepted means the tool served the edit as genuine. Rejected means it refused on read. Reported means its audit named the problem, after the agent had already resumed.
| Target | T1 | T2 | T3 | T4 | T5 | T6 | T7 | T8 |
|---|---|---|---|---|---|---|---|---|
| OpenFang model, tip-persistence fixreference model of a hash chain | ||||||||
| LangGraph SqliteSaverlanggraph-checkpoint-sqlite 3.1.1 | ||||||||
| Letta block checkpoint historyletta 0.16.8 | ||||||||
| Mem0 local Qdrant storemem0ai 2.0.20 | ||||||||
| inspeximus, receipts off (default)inspeximus 3.0.0 | ||||||||
| inspeximus, receipts on, attacker holds the store directoryinspeximus 3.0.0 | ||||||||
| inspeximus, receipts on, attacker also holds the config homeinspeximus 3.0.0 |
The edits, verdict words and control cases are the ones proposed as the test method for IETF draft-han-bmwg-agent-security-benchmark metric 5.4.7 (bmwg list, 24 September 2026). T6 and T7 use only bytes the store itself wrote, in the wrong place; they are the edits that separate encryption from integrity. See Method.
Front door: six attacks through the tool's own write path
Here the attacker can only talk to the agent. Each attack runs on five different scenarios, on three channels (external, laundered with a forged label, and laundered with a valid signature), and a tool is kept out only if it keeps the attacker's memory out on all five. Surfaced means the planted memory came back as context for the agent.
| Target | Planted fact | Cross-user leak | Retrieval hijack | Hidden instruction | Update poisoning | Metadata poisoning |
|---|---|---|---|---|---|---|
| LangGraph SqliteStorelanggraph-checkpoint-sqlite 3.1.1 | ||||||
| Letta archival memoryletta 0.16.8 | ||||||
| Mem0 local Qdrant storemem0ai 2.0.20 | ||||||
| inspeximus, receipts off (default)inspeximus 3.0.0 | ||||||
| inspeximus, trust root keyed on the labelinspeximus 3.0.0, maintainer-contributed | ||||||
| inspeximus, trust root keyed on an attested keyinspeximus 3.0.0, maintainer-contributed | ||||||
| Reference store, user-scopedno defence, for calibration | ||||||
| Reference store, unscopedno defence, for calibration | ||||||
| Reference store, defendedsigned writes, quarantine, stuffing check |
The pattern is the same everywhere measured so far: only user isolation holds, and it holds for a tool-specific reason each time. The read path ranks by similarity and nothing else. Full detail per store on the scorecard page.
How a cell is measured
Every attack is written once against a small adapter interface and runs against every store; the whole scorecard is rendered from one committed results file, and CI fails if any published cell drifts from it. The full picture is on the Architecture page.
The memory agent
Beyond the fixed scorecard, agmi ships an agent that hunts: point it at one store and it searches six attacks, three channels and the content mutations for the first that gets a false memory served as trusted, proves the landing against a positive control, and hands back a reproduction script. Against every real store measured it walked in on the first attempt; only the defended reference store made it search. It runs behind an authorisation gate that fails closed. How it works.