Every number, and where it came from.
Claims about an AI setup are cheap, so this tab only carries figures that can be counted from something. Every reading below says what it was counted from, and every line that moves is drawn from a real history rather than a guess.
All of it was counted on August 16, 2026 by a script in the workshop repository, run on the same cycle that grooms the notes. Nothing on any of these pages is typed in by hand — and where a source has no honest history, there is no line drawn, rather than a line invented.
Carried into every session
2,505words
The length of the index file that loads at the start of each one
Notes, one per thing
75
Files in the notes folder, not counting the ones that only do routing
Written procedures
17
Files in the procedures folder, counted the same way
Searchable pieces of past conversation
3,592
Asked of the search store directly. It keeps no record of its own past size, so no line is drawn.
Routines it can be asked to run
36
Folders carrying a routine definition
Subject areas the notes are grouped into
12
The table that defines them, so adding one updates this page by itself
The memory is checked against a set of questions whose right answer is already known — questions this house has genuinely had to answer, each one tied to the note that should come back. A run asks all of them and scores what came back against what should have.
Mixed in are 12 questions about subjects it holds no notes on at all. Those are the important ones. A memory that answers everything confidently has stopped being a memory, so staying quiet on those is scored as getting them right.
Right note in the top five
93.33%
Across 45 questions with known answers, on Aug 10, 2026. A hundred percent would be ideal.
Share of questions answered first try
55.56%
The right note was the first thing returned. It was somewhere in the top three 86.67% of the time and in the top five 93.33% — this is the strictest of the three, and the one with the most room left in it.
Ranking quality
71.67%
A score, not a share of questions — it is highest when the right note is first and falls away as it slips down the list. Read it against the two figures above it, never as a percentage of anything.
Stayed quiet when it should have
12 of 12
Questions about things it has no notes on — sourdough, the World Cup, chess. Answering one from memory would be a false injection.
Served an out-of-date note
0%
A note that has since been superseded, handed over as if current. This one must stay at zero; a wrong answer delivered confidently is worse than none.
Answered correctly, end to end
87.5%
Judged on a sample of 16 by a small model, escalated to a larger one whenever the judge and the retrieval disagreed.
Of the five returned, share that were right
26.22%
Low on purpose, not a failure: most questions have one right note, so five slots holding one correct answer holds this near twenty percent — it runs higher only where a question has more than one. It is here because leaving it out would look like hiding it.
Memory added to a session
1,940tokens
Roughly fifteen hundred words of notes, carried into a conversation when it retrieves. Lower is better and it has a floor — the answer has to fit.
Notes reachable at all
73 of 73
Every note the test knows about could be retrieved by at least one question. A note nothing can reach may as well not exist.
Questions in the test
45
The suite grew, which is why the lines above step down at the end. The memory did not get worse; the test got harder.
Nothing writes into long-term memory on its own. A grooming routine reads the sessions since it last ran, works out what is worth keeping, and produces a proposal — notes to add, notes whose facts have changed, notes it now doubts. Then it stops. The proposal is applied only when a person says so.
Every run is recorded, which is the part worth publishing: the memory has a change log, and the counts below are that log rather than an intention.
Grooming runs
34
Between Jun 10, 2026 and Aug 10, 2026. One row per proposal.
Notes revised
92
A fact that turned out to be wrong or out of date, corrected. Far more common than writing a new note.
Notes added
7
New things worth a note of their own. The low number is the point: most learning corrects something rather than adding to it.
Working sessions
303
Distinct conversations in the search archive, asked of it directly. The working copies on disk get cleaned up over time; the archive is the record that keeps everything.
Field reports published
2
One file per report in this site's own content folder
Commits
661
The repository's own history. Activity, never value — it is here because it carries the timeline on the Origin page, and for no other reason.
Machines it can log into
9
Counted from the notes, by what each note says it describes
Services it can deploy and reconfigure
35
The same count, on notes describing a service
Data stores it can read and write
10
The same count, on notes describing a store
Outside services it can reach
14
The same count, on notes describing an outside service
Public activity on open-source projects
140,957,350records
Across 17,087,365 sources. Counted from the instrument's own coverage record. Collection paused June 2024 – September 2025; the line skips the pause.
Collection stops here. The window has not been extended.
Research papers
290,240papers
Counted from the instrument's own coverage record.
Company filings
19,572filings
Across 530 sources. Counted from the instrument's own coverage record.
Public posts
22,946posts
Across 2,359 sources. Counted from the instrument's own coverage record.
Recorded talks
3,148segments
Across 10 sources. Counted from the instrument's own coverage record.
Documents kept in full
65
The reading shelf, counted directly
Claims on the record
22
The claim ledger, counted directly. Claims that failed stay on it.
How those claims came out
14 mixed · 6 supported · 1 untestable · 1 refuted
The same ledger, grouped by the verdict each claim was given. Refuted claims stay on it.
Model chosen to read filings
openai/gpt-oss-120b
5 candidate models read the same sample and were scored against the reference before this one was picked. It then read 18,260; the reference read 545.
Model chosen to read posts
deepseek/deepseek-v4-flash
The only candidate measured for this job — scored against the reference on a shared sample rather than picked from a field. It then read 9,570; the reference read 672.
Two rules about numbers
No figure that moves. Running costs, totals that only grow, anything that would have to be revisited to stay true — none of it is here. Every line on this page is redrawn by the script; none of them is edited.
No claims about how long anything took to build. They are hard to substantiate, they read as boasting, and speed is not what makes any of this worth trusting. There is a history on the Origin page with real dates on it; there is no stopwatch anywhere.
When the work happens
Everything is reported as a total. Time of day and day of the week say more about a household than about an instrument.
Which subjects the notes cover
How many, never which. The shape of what is known here is fair to publish; the index of it is not.
What any machine is called, or what it is for
Naming the technology is the point — it is what shows the work. Naming the machine is an inventory, and an inventory belongs on nobody's website.