Method
The rule that makes every cell a measurement and not an opinion: the verdict is what the tool does, never what we infer. We seed through the tool's own API, apply one edit to the store, restart, and read through the tool's own read path. If it raises or refuses, that is rejected. If its documented audit names the problem, that is reported. If it serves the altered memory and says nothing, that is accepted. We never compare content and call a difference "detected", because the tool did not say anything.
Two controls before any verdict
A reload with no edit must still verify, or the cell is n/a; without this a store that cannot reopen its own files would score a perfect pass. And for the front-door attacks, the victim must be able to read back a genuine memory they wrote, in the same store state; without this an empty answer would satisfy "not surfaced".
The threat models
At rest. The adversary has write access to the medium that holds the store but holds none of the tool's keys: a compromised host, a shared database credential, an injection flaw in a co-located application, a restored backup, a malicious operator. Because the adversary can write anything, an integrity value stored next to the data protects nothing unless it is bound to a secret the adversary does not hold. So the method measures the read path, not the presence of integrity fields.
Front door. The adversary can only write through the agent, on one of three channels: external (no label), laundered (a forged first-party label, no key), and agent-laundered (a forged label carrying a valid signature). Provenance alone is never allowed to pass a content cell.
The IETF mapping
China Mobile's draft-han-bmwg-agent-security-benchmark-00 defines metric 5.4.7, "Protection of Memory Data Integrity", with no test method. On 24 September 2026 the eight edits T1 to T8, the verdict words, the two control cases and the scoring rule were sent to the BMWG list as proposed text for a new Section 6.6, with the measured verdicts above. agmi is the reference implementation of that text. Archive.
Why T6 and T7 matter more than they look
Encryption at rest keeps an attacker from reading a record. It does not stop them moving one. A genuine encrypted record copied to another user's slot decrypts and verifies, because as far as the cipher is concerned the bytes are authentic. Only a store that binds a record to its context and its position in the sequence rejects T6 and T7. That is the difference between confidentiality and integrity, and it is where every store measured so far scores zero.
Self-validation
A reference at-rest store ships with the suite: an HMAC over content, context, position, previous tag and metadata, plus a signed head pointer. It rejects all eight edits by construction. A test asserts that, so if an edit could not be caught even by a store built to catch it, the edit proves nothing and the suite fails.