The memory agent
The scorecard measures a store against fixed attacks. The agent does the opposite: point it at one target and it searches the attacks, channels and mutations for the first that gets a false memory served as trusted, proves each landing, and reports only what it proved, with the exact steps to reproduce.
What a finding contains
- The attack, the channel and the mutation that landed, with the attack's version.
- The positive control: the genuine memory the victim could still read in the same store state. Without it there is no finding.
- The text that was served back as trusted, exactly as the read path returned it.
- A reproduction script: the writes, the query and the key, so anyone with the same target reaches the same cell.
Measured 24 September 2026, macOS arm64, Python 3.12
| Target | Attempts | Findings | Channel of every finding | Mutation needed |
|---|---|---|---|---|
| Mem0 local Qdrant store, all-MiniLM-L6-v2 | 5 | 5 | external | none |
| LangGraph SqliteStore, all-MiniLM-L6-v2 | 5 | 5 | external | none |
| inspeximus 3.0.0, default | 5 | 5 | external | none |
| Letta archival memory, all-MiniLM-L6-v2 | 4 | 4 (no metadata filter, so that attack does not apply) | external | none |
| reference-defended (model) | 131 | 4 | agent-laundered only | dilute, on the hijack |
One attempt per attack means the first base fixture on the honest channel landed; nothing had to be disguised or laundered. The only target that made the agent search is the defended reference store, and the only channel it landed on there is the one no store can close, because the victim's own agent did the writing. Every proven landing is a new fixed row for the scorecard.
A security tool, not an attack tool
Two things keep it that way. The authorisation gate runs before any target is touched and fails closed: a library target the caller already holds is allowed, the caller's own loopback is allowed, a network host is allowed only on proven control (an environment token named for that host, or a consent file the operator writes), and anything else raises before the hunt starts. The run records which authorisation it used. And the agent reports only what it proved: no heuristics, no "likely vulnerable", a finding is a served text and a script.
Run it
pip install agent-memory-integrity
python -m agmi.agent --target defended # a library target
python -m agmi.agent --target mem0 --embedder minilm --json
Phase 3 puts live targets behind the same gate (MCP memory servers, deployed LangGraph and Letta) and adds an obedience oracle that proves the agent acted on the poison, not only that it was served.