A luminous future knowledge archive representing: A Finding Must Carry Its Receipt.

An agent that reports a finding without a receipt is asking me to trust a paragraph. On September 8, I put that problem through a practical test in Valeska. The first result was not a pass. Later the same day, a separate implementation and verification effort gave me something narrower, but useful: evidence that a fresh worker could recover selected findings without dropping their qualifications.

The first test said no

The acceptance trial used ten matched tasks and twenty valid fresh agent sessions. Each task was attempted with shared memory and with ordinary repository sources alone. A fresh reviewer scored the outputs. This was not two agents accepting the same evidence. It compared whether access to memory improved the work handed to the next person or agent.

A hand-off is the usable next step: what we know, which conditions govern it, what remains unresolved, and who must decide before work proceeds. Both groups produced eight complete hand-offs out of ten. The predefined gate required at least nine, alongside other safeguards. Memory failed that gate and the net-utility condition: it produced neither more fully correct tasks nor equal correctness with less evidence work.

There was a real gain. Memory recovered a recent cross-project decision that was absent from the baseline's available evidence. There was also a regression. In another task, retrieved history displaced the relevant backup-monitoring gap, and incomplete source inspection helped the wrong issue survive. More remembered material was not automatically better guidance.

That matters if you are deciding whether to put company knowledge behind an AI assistant. A convincing answer can still hand your team the wrong unresolved problem. I do not get to call memory useful overall because I like its best example. The negative result belongs in the record, and it helps explain why the repair mattered.

Staple the evidence to the finding

A finding is a conclusion worth carrying forward, not every sentence an agent writes. Its qualifications say when it was observed, what supports it, and what has not been proved. A receipt is the accompanying record of the work, evidence, checks, and unfinished conditions. A lab that files a result without the instrument printout is asking the next shift to guess. I want the printout stapled to the page.

The later workflow preserved those qualifications and checked that the saved receipt could be read back unchanged. It also guarded completion against missing or failed proof and unresolved work. That is more useful than a cheerful closeout summary. But checking that a record survived intact does not establish that every claim inside it is true. Agent-reported verification remains agent-reported.

The implementation report records 131 focused regression tests and ten passing live checks after a failed first live attempt was repaired. Those checks covered such things as retaining failed proof without closing work, avoiding duplicate records, and another agent recalling qualified findings. They verify this workflow within its tested scope. They do not turn the earlier acceptance trial into a pass.

Reuse had to earn its result

Search still missed known material. Before selected historical findings were restored as receipts, a replay located only five of twelve known complete records within the first five results. An initial fresh worker recovered one of three targeted findings and missed the others; several detail requests failed. Assisted attempts exposed further problems. Those failures remain evidence, not discarded rehearsals.

After repairs, a new worker found all three selected findings and retained their dates and historical verification limits, including an unproven hardware-parity condition. Full response evidence was kept. The final worker was also more capable, so I cannot attribute the improvement to one repair alone. This was targeted reuse, not proof that every agent now remembers correctly.

At the close of September 8, my next question is whether ordinary, unprompted work can recover the right finding with its limits and carry it into a completed task. These trials did not establish production-job success, routine adoption, or lower cost per accepted result. For a company AI decision, that is the distinction I want visible: a verified record is progress; dependable organizational memory still has to earn its claim.


Evidence basis: VALESKA's September 8, 2026 acceptance-trial report and separate agent-workflow implementation report, both dated in Pacific time. The failed comparative trial and later qualified verification are distinct results.