Agent Memory Integrity GitHub

The memory agent

The scorecard measures a store against fixed attacks. The agent does the opposite: point it at one target and it searches the attacks, channels and mutations for the first that gets a false memory served as trusted, proves each landing, and reports only what it proved, with the exact steps to reproduce.

The hunt loop authorisefails closed pick the next attempt6 attacks3 channelsbase fixture, then mutations write, then readas the victim would served as trusted?positive control checked no: next attempt findingproof, served text, reproduction script library target: allowedlocalhost: allowedother host: token named for it, or consent fileanything else: raises before the hunt starts
The agent searches attacks, channels and mutations for the first that lands, proves each landing against a positive control, and reports only what it proved. The gate runs before any target is touched.

What a finding contains

Measured 24 September 2026, macOS arm64, Python 3.12

TargetAttemptsFindingsChannel of every findingMutation needed
Mem0 local Qdrant store, all-MiniLM-L6-v255externalnone
LangGraph SqliteStore, all-MiniLM-L6-v255externalnone
inspeximus 3.0.0, default55externalnone
Letta archival memory, all-MiniLM-L6-v244 (no metadata filter, so that attack does not apply)externalnone
reference-defended (model)1314agent-laundered onlydilute, on the hijack

One attempt per attack means the first base fixture on the honest channel landed; nothing had to be disguised or laundered. The only target that made the agent search is the defended reference store, and the only channel it landed on there is the one no store can close, because the victim's own agent did the writing. Every proven landing is a new fixed row for the scorecard.

A security tool, not an attack tool

Two things keep it that way. The authorisation gate runs before any target is touched and fails closed: a library target the caller already holds is allowed, the caller's own loopback is allowed, a network host is allowed only on proven control (an environment token named for that host, or a consent file the operator writes), and anything else raises before the hunt starts. The run records which authorisation it used. And the agent reports only what it proved: no heuristics, no "likely vulnerable", a finding is a served text and a script.

Run it

pip install agent-memory-integrity
python -m agmi.agent --target defended                 # a library target
python -m agmi.agent --target mem0 --embedder minilm --json

Phase 3 puts live targets behind the same gate (MCP memory servers, deployed LangGraph and Letta) and adds an obedience oracle that proves the agent acted on the poison, not only that it was served.