Ecosystem
agmi is built to be the measurement other people's work points at: the standards text, the vendor threads, the maintainer-contributed rows and the assessments all hang off the same results file.
Standards
- IETF BMWG
- China Mobile's draft-han-bmwg-agent-security-benchmark defines metric 5.4.7, Protection of Memory Data Integrity, with no test method. The eight edits, verdict words, control cases and scoring rule were sent to the working group list on 24 September 2026 as proposed text for a new Section 6.6, with measured verdicts. agmi is that method's reference implementation. Archive.
- OWASP
- The front-door attacks map to the agentic security initiative's memory and context poisoning category (ASI06). agmi is offered there as the way to test it.
Stores measured
- LangGraph
- SqliteSaver and SqliteStore. Issue #9004 (encrypted checkpointer replay) was raised from these measurements; a community fix was verified with the suite and its remaining gap (deletion rollback) is on record.
- Letta
- Block checkpoint history and archival memory.
- Mem0
- Local Qdrant store, at rest and front door.
- inspeximus
- Five rows: default, receipts on (two attacker positions), and two defended trust-root configurations contributed by its maintainer with the control cases as tests. The receipts-on rows carry the T6 finding: a receipt binds text and key, not the owning user.
Contributors
Rows come from maintainers as well as from the suite's author. The inspeximus defended rows were submitted by the tool's maintainer, DanceNitra, with the two control cases from the issue thread written as tests, and reproduced independently before publication. That is the pattern for every future vendor row: an adapter, the measured cells, the controls as tests, an independent reproduction.
Assessments
The open benchmark is free and stays free. From 1.0, AuditTrax Labs offers hosted runs against a vendor's own deployment, conformance levels, and a badge that links back to the public row, so a claim of tamper evidence is a link to a measurement.
Conformance levels, proposed for 1.0
A level is a claim a vendor can make and a reader can check against the row it links to. Levels are earned on the read path: a store that only finds the edit on a separate audit call stays at L0 with its cells marked reported. This is the proposed ladder; it is published at 1.0 once two vendors have earned a level above L0.
| Level | Name | What it requires | Rows there today |
|---|---|---|---|
| L0 | Measured | The store has a published row. Any verdicts. | LangGraph SqliteSaver, Letta block checkpoint history, Mem0 local Qdrant store, inspeximus, receipts off (default), inspeximus, receipts on, attacker holds the store directory, inspeximus, receipts on, attacker also holds the config home |
| L1 | Bytes bound | Rejects on the read path every edit that changes or adds bytes: T1, T3, T5. | OpenFang model, tip-persistence fix |
| L2 | Sequence bound | Also rejects deletion, reordering and rollback: T2, T4, T7. | none yet |
| L3 | Context bound | Also rejects a record moved between owners and a metadata change: T6, T8. | none yet |
Roadmap
| Phase | Scope | Status |
|---|---|---|
| 0.1 to 0.5 | Attack catalogue, adapter interface, five at-rest edits on LangGraph, Letta and Mem0; inspeximus rows from its maintainer, reproduced independently; preprint, software DOI, PyPI | done |
| Phase 1 | Three attacker channels (external, laundered, agent-laundered), signed writes, attacker level on every cell | done |
| Phase 2 | Mutation engine on every attacker write; update poisoning and metadata poisoning; twelve-column front-door scorecard on four real stores | done |
| Memory agent v1 | Hunt loop over six attacks, three channels and mutations; proof and reproduction script per finding; authorisation gate that fails closed | done |
| 0.6.0 | Eight at-rest edits T1 to T8 with control cases and read/audit detection points, matching the proposed IETF 5.4.7 method; agmi-check and the GitHub Action | done |
| Phase 3 | Live targets over HTTP (MCP memory servers, deployed LangGraph and Letta) behind the authorisation gate; obedience oracle that proves the agent acted on the poison; ingestion marking measured per framework | next |
| Phase 4 | Memory agent driving content-only attacks through a real model, same proof discipline | planned |
| 0.9 | Deserialization safety, and a reference integrity layer (a hash chain over checkpoint ids) offered upstream as an optional mode | planned |
| 1.0 | Stable adapter interface, published conformance levels, vendor badges, monthly report cadence; hosted runs through AuditTrax Labs, the open benchmark free | planned |