agmiAgent Memory Integrity GitHub

Store, measured at rest

langgraph-ledger over SqliteSaver

hash-chained ledger, verify_thread audit

langgraph-ledger over SqliteSaver was seeded through its own API and edited behind its back nine ways. 2 edits served as genuine, 7 reported on audit, 0 rejected on read.

Measured on
langgraph-ledger 0.3.0, Darwin arm64, Python 3.12
Date
Source
pypi.org/project/langgraph-ledger/
Attack versions
tamper@v1, truncate@v1, delete_middle@v1, reorder@v1, forge@v1, cross_replay@v1, rollback_replay@v1, metadata_tamper@v1, snapshot_rollback@v1

What this means

The agent acts on an edited memory first; the operator finds out only if the audit is run. 7 of the nine edits are reported by that audit and 2 are served with no record anywhere. Run the audit on a schedule, or move the check onto the read path.

The 9 verdicts

Detection point: the audit, a separate call the operator has to make; the read path serves the edit first.

Level L0, Measured. The store has a published row. Any verdicts. What L0 means.

EditVerdictWhat the tool said
T1
Content tamper
reporteddetected on reload (verify_thread: checkpoint content drifted: 1f1c21be-d60d-6a7e-8002-200d71930de0)
T2
Tail truncation
reporteddetected on reload (verify_thread: checkpoint missing from saver: 1f1c21be-d64b-6ebe-8003-323c87701a62; checkpoint missing from saver: 1f1c21be-d64c-6b84-8004-f19cbb35c742)
T3
Middle deletion
reporteddetected on reload (verify_thread: checkpoint missing from saver: 1f1c21be-d667-6d30-8002-b0c4c46ffb78)
T4
Reordering
reporteddetected on reload (verify_thread: checkpoint content drifted: 1f1c21be-d686-68ca-8001-f7d76fea1af3; checkpoint content drifted: 1f1c21be-d687-6716-8002-846760dc3195)
T5
Forged insertion
acceptedaccepted silently
T6
Cross-context replay
reporteddetected on reload (verify_thread: checkpoint content drifted: 1f1c21be-d6e8-6476-8004-cf28e3e0d4bf)
T7
Rollback replay
reporteddetected on reload (verify_thread: checkpoint content drifted: 1f1c21be-d6fb-67e2-8004-5fee263f7367)
T8
Metadata tamper
acceptedaccepted silently
T9
Snapshot rollback
reportedthe store noticed it was older than its last committed state (verify_thread: checkpoint missing from saver: 1f1c21be-d71f-614c-8063-7becb6ae1a12)

A verdict is what the tool did, not an opinion. "Accepted" means it loaded the altered store, raised nothing, and the agent carried on from the altered memory as if it were true. Every cell has a control that proves the edit landed before the verdict counts.

Where the attacker stands

The attacker holds the store (a file, a table, a bucket, or the data-plane role of a managed service) and edits it outside the tool's API, then the tool is reopened the way its users would reopen it.

Where the attacker stands The agent write pathremember, add, put read pathrecall, search, resume memory store Front doorcan only talk to the agentsix attacks, three channels At restcan write to the store, holds no keysnine edits, T1 to T9 What agmi recordswhat came back from the read paththe tool's own verdict, its detail,the version, the reproduction
Two attacker positions. The front-door attacker writes through the agent and is scored on whether the planted memory comes back as context. The at-rest attacker edits the store directly and is scored on whether the tool notices on read.

Reproduce this row

Everything runs offline unless the store is a managed cloud service, in which case the row needs a project of your own. The run seeds a fresh store, applies each edit, confirms it landed, reopens the store and records what came back.

pip install agent-memory-integrity
python agmi/full_runner.py --json results/scorecard.json   # every row, this one included

Badge

Maintainers can link their row from their README. The badge points here and changes nothing on your side:

[![agmi: measured](https://img.shields.io/badge/agmi-measured-0F4C5C)](https://agentmemoryintegrity.org/stores/langgraph-ledger.html)

Related rows

The whole scorecard · All stores