●  project in progress updated 4 October 2026
The promise

A memory for AI agents that never serves as true a memory its source does not support.

Every memory is tied to its source. A local judge decides whether the source really supports it. An unsupported fact is quarantined. With a receipt: who wrote the source, when, the content fingerprint, who decided.

Not for sale. The code is public. This page says what works, what does not and with which numbers, and it changes when the numbers change.

What the first version did, and what it did not keep

Version 0.7, 4,705 commits from July to September 2026, is public at github.com/aureliocpr-ctrl/verimem. Measured by us, on our machine.

The promiseWhat we measuredWhere
A fact the source does not state is quarantinedAn invented detail without numbers was admitted 8 times out of 10 in Italian and 9 out of 10 in English; with a number inside it was almost always stopped0.7 README, line 37
No memory gets lost21.1% of memories had never been served: 3,828 out of 18,147, as of 13 September 2026register of 13 September
Every memory carries its sourceThe store kept a fingerprint, not the source: the full source can be recovered for 628 memories out of 9,245, 6.8%register of 3 October

That is why 0.7 has been stopped since 30 September 2026. We do not sell it and we do not recommend it. github.com/aureliocpr-ctrl/verimem ↗

The new core

On 2 October 2026 an instance without the context of the old version rewrote the core from scratch: github.com/aureliocpr-ctrl/verimem-next, Apache-2.0, 47 commits, 190 tests, 3,165 lines. One write path. No training: the judge is an off-the-shelf model.

  • On the real memories of the pilot, the base judge without retraining scores 0.962 AUROC against 0.789 for the judge we had retrained: retraining on labels given by a model had made it worse.
  • On an Italian benchmark with human labels (EVALITA 2009, 800 pairs), the best off-the-shelf judges score 0.880, 0.870 and 0.841: none is perfect, and we say so.

github.com/aureliocpr-ctrl/verimem-next ↗

What we found while studying

The machine speaks in your voice.

In agent transcripts, 85.6% of the blocks marked "user" is text injected by the machine. Whoever reads without a filter remembers the machine as if it were you. None of the tools we took apart filters it.

Coherence is not truth.

If the agent picks the source, the judge checks that the memory matches that text, not that the text is true. That is why the receipt will always say whether the source was read by the system, pointed to, declared, or absent.

Who speaks, whose fact it is, what kind of line it is.

"My dog is called Aron" said by you is a fact about you; said by the assistant, quoted, asked or hoped, it is not. Writing the voice into the text before judging, errors across five judges and four languages drop from two thirds to almost zero. The kind of line, question, plan, outburst, quote, is the next experiment.

Repeating does not make it true.

A fact said three times on three days is "said three times", with dates and lines next to it, never "true". Only the person's voice counts, one vote per day, and the machine's echo never counts.

The lab: how we decide whether the road is right

Six questions with criteria written beforehand and dates.

  1. Does the judge hold on real memories? If on 523 human-labelled cases it does not keep nine true memories out of ten while letting at most one false in ten through, and the cures outside the model do not save it, the judge leaves the product. by 23 October 2026
  2. Do the cures outside the model hold on real data? A cure that costs more than 5% of true memories or holds only on synthetic data leaves. by 23 October 2026
  3. Does the source-memory link exist in other memory stores? Yes in three out of five: where the source remains, the backward report is possible. decided
  4. Is the cost sustainable? A light judge under one gigabyte and under two tenths of a second per fact, or a resident service. by 23 October 2026
  5. Does anyone want it? Forty conversations. With fewer than three replies, it stays our own tool. by 1 December 2026
  6. Is the new core a base? Yes: nineteen defects found, none structural. decided

The first product that can close is a report on the memories an agent already has: for each memory, who wrote the source, from which source, how old it is, whether the source changed afterwards, whether numbers and dates match, whether the source really says it, and who decided. The first layer needs no model at all.

Write to me

If you have an agent with a memory and want to know how far you can trust what it remembers, write to me.

aureliocpr@gmail.com github.com/aureliocpr-ctrl/verimem-next ↗ Who I am and how I work: capriello.com ↗