The Lab

What it remembers, and what it refuses to trust.

A model forgets everything the moment a conversation ends. Three layers sit underneath this one so it doesn’t — and a set of rules sits over those, because a memory that is confidently wrong is worse than no memory at all.


A sessionREADSAlways loaded2,505 wordsevery sessionLooked up128 fileson the subjectSearched3,592 pieceson demandautomaticallyonce approved
The layers differ in one respect that decides everything else: how often each one is present. What is always carried is paid for in every conversation, so it stays short and holds pointers. The two paths back are not the same act — the archive fills itself, while a durable note is only ever written after a person approves it.

The three layers


What it goes and looks up

Holds
How each thing genuinely works — the connection details, the settings that matter, and the mistakes already made against it
Loads
Only on the subject at hand. The notes are sorted into 12 subject areas, and entering a subject pulls in that neighborhood rather than the whole library
The useful half
A note records the scar as well as the fact. Knowing a port is one thing; knowing that changing it silently broke something last time is what stops the mistake from being made twice
Confidence
Notes are marked for how sure they are. Anything marked uncertain or aging is verified before it is relied on, rather than quoted with confidence it has not earned
Three layers. Click one for its detail. Everything above the archive was written deliberately; the archive is the only layer that fills itself.

What a conversation does

THIS SESSIONLoadWorkGateWriteTHE NEXT ONELoadwhat it learned
A session does not start empty and does not end when the answer arrives. It ends by leaving something behind — which is the only reason the next one starts further along. The gate is where anything irreversible stops and waits for a person.

Three rules that keep it honest

Building a memory is the easy half. The hard half is making sure it does not become a machine for repeating things that were never true, or that stopped being true months ago. Three rules do that work.

The archive is a record of what was said, not a claim about what is true

Anything retrieved from it is quoted with its date and checked against the live system before it is acted on. A confident sentence from June is evidence about June, and nothing more. If it disagrees with what is actually running, the transcript loses.

Notes carry their own confidence

A note written from direct observation and a note inferred from context are not the same thing, and they are not marked the same way. Anything uncertain or aging gets verified before it is relied on for something that matters, rather than being quoted with a certainty it never earned.

Nothing writes into memory unsupervised

The grooming routine reads recent sessions, drafts what it thinks is worth keeping, and then stops and asks. It has run 34 times, and every run is on the record: what it added, what it corrected, what it flagged as no longer trustworthy. A system that edits its own long-term memory without review will eventually record something it invented, and from then on repeat it with perfect confidence forever. The approval step is the entire defense against that, and it is not automated away.

None of these are technical problems. All three are decisions about what the machine is not allowed to do on its own.

Whether it works

All of the above is a claim, and claims are cheap. So the memory is checked against a set of questions whose right answer is already known — real questions this house has had to answer, each one tied to the note that should come back. A run asks all of them and scores what actually came back.

Mixed in are 12 questions about subjects it holds no notes on at all: sourdough, the World Cup, chess. Those are the important ones. A memory that has something to say about everything is not a memory, it is a liability, and staying quiet on those is scored as getting them right.

93.33%
of the time the right note came back in the top five
55.56%
of the time it came back first
12/12
questions about things it has no notes on, answered by staying quiet
87.5%
judged correct end to end, on a sample of 16
The most recent of 8 runs, on 45 questions. Ranking the right note first is the number with the most room left in it, and it is the one worth watching — being in a list of five is easier than being the answer.

One honest note about those figures. Six straight runs scored a flat hundred percent on the top-five measure, and then it fell. Nothing broke: the test grew by half, the new questions were harder, and the score followed. A number that only ever improves is usually measuring the wrong thing, or measuring it too gently.