What it remembers, and what it refuses to trust.
A model forgets everything the moment a conversation ends. Three layers sit underneath this one so it doesn’t — and a set of rules sits over those, because a memory that is confidently wrong is worse than no memory at all.
The three layers
What it goes and looks up
- Holds
- How each thing genuinely works — the connection details, the settings that matter, and the mistakes already made against it
- Loads
- Only on the subject at hand. The notes are sorted into 12 subject areas, and entering a subject pulls in that neighborhood rather than the whole library
- The useful half
- A note records the scar as well as the fact. Knowing a port is one thing; knowing that changing it silently broke something last time is what stops the mistake from being made twice
- Confidence
- Notes are marked for how sure they are. Anything marked uncertain or aging is verified before it is relied on, rather than quoted with confidence it has not earned
What a conversation does
Three rules that keep it honest
Building a memory is the easy half. The hard half is making sure it does not become a machine for repeating things that were never true, or that stopped being true months ago. Three rules do that work.
The archive is a record of what was said, not a claim about what is true
Anything retrieved from it is quoted with its date and checked against the live system before it is acted on. A confident sentence from June is evidence about June, and nothing more. If it disagrees with what is actually running, the transcript loses.
Notes carry their own confidence
A note written from direct observation and a note inferred from context are not the same thing, and they are not marked the same way. Anything uncertain or aging gets verified before it is relied on for something that matters, rather than being quoted with a certainty it never earned.
Nothing writes into memory unsupervised
The grooming routine reads recent sessions, drafts what it thinks is worth keeping, and then stops and asks. It has run 34 times, and every run is on the record: what it added, what it corrected, what it flagged as no longer trustworthy. A system that edits its own long-term memory without review will eventually record something it invented, and from then on repeat it with perfect confidence forever. The approval step is the entire defense against that, and it is not automated away.
Whether it works
All of the above is a claim, and claims are cheap. So the memory is checked against a set of questions whose right answer is already known — real questions this house has had to answer, each one tied to the note that should come back. A run asks all of them and scores what actually came back.
Mixed in are 12 questions about subjects it holds no notes on at all: sourdough, the World Cup, chess. Those are the important ones. A memory that has something to say about everything is not a memory, it is a liability, and staying quiet on those is scored as getting them right.
- 93.33%
- of the time the right note came back in the top five
- 55.56%
- of the time it came back first
- 12/12
- questions about things it has no notes on, answered by staying quiet
- 87.5%
- judged correct end to end, on a sample of 16
One honest note about those figures. Six straight runs scored a flat hundred percent on the top-five measure, and then it fell. Nothing broke: the test grew by half, the new questions were harder, and the score followed. A number that only ever improves is usually measuring the wrong thing, or measuring it too gently.