Every number, and where it came from.
Claims about an AI setup are cheap, so this tab only carries figures that can be counted from something. Every reading below says what it was counted from, and every line that moves is drawn from a real history rather than a guess.
All of it was counted on September 15, 2026 by a script in the workshop repository, run every time the notes are groomed, a session is closed out, or the memory is tuned. Nothing on any of these pages is typed in by hand — and where a source has no honest history, there is no line drawn, rather than a line invented.
The first five panels mirror the dashboard the agent’s own control panel shows its author, panel for panel and in the same order. What that dashboard shows and this page does not: the name of any machine, the name of any subject, and any bill.
Last groomed
9days ago
Nothing writes into long-term memory on its own. A grooming routine reads the sessions since it last ran, works out what is worth keeping, and produces a proposal — notes to add, notes whose facts have changed, notes it now doubts. Then it stops. The proposal is applied only when a person says so, and every run is recorded: the memory has a change log, and the counts here are that log rather than an intention.
Grooming runs
36
Between Jun 10, 2026 and Sep 6, 2026. One row per proposal.
Notes revised
120
A fact that turned out to be wrong or out of date, corrected. Far more common than writing a new note.
Notes added
7
New things worth a note of their own. The low number is the point: most learning corrects something rather than adding to it.
Procedures touched
12
Written procedures added or revised by a grooming run, over the same window.
Two other models — Gemini and GPT, rented the same way the main one is — read the same proposal before it is applied. On each item a reviewer either agrees with the groomer or files a disagreement; a disagreement whose fix is better goes into memory, and one that is not is overruled. A run where both simply agreed is the healthy default, not a failing score.
Gemini 2.5 Pro
Agreed with the groomer on all 13 items
- 12agreed with the groomer
- 0disagreed · the groomer's version kept
- 0disagreed · the reviewer's fix accepted
GPT-4.1
Agreed with the groomer on all 13 items
- 13agreed with the groomer
- 0disagreed · the groomer's version kept
- 0disagreed · the reviewer's fix accepted
The memory is checked against a set of questions whose right answer is already known — questions this house has genuinely had to answer, each one tied to the note that should come back. A run asks all of them and scores what came back against what should have.
Mixed in are 12 questions about subjects it holds no notes on at all. Those are the important ones. A memory that answers everything confidently has stopped being a memory, so staying quiet on those is scored as getting them right.
Every one of the 124 notes and procedures can be reached by at least one question in the test. A note nothing can reach cannot miss, so this line comes before any score.
Right note in the top five
100%
ideal 100%
Across 45 questions with known answers, on Sep 6, 2026.
Ranking quality
93.22%
a score, not a share · first 88.89% of the time
Highest when the right note is first and falling away as it slips down the list. Read it against the two tiles beside it, never as a percentage of anything.
Of the five returned, share that were right
41.15%
low by design · about one right answer per question
Most questions have one right note, so five slots holding one correct answer keeps this near twenty percent — it runs higher only where a question has more than one. It is here because leaving it out would look like hiding it.
Served an out-of-date note
0%
must stay 0%
A note that has since been superseded, handed over as if current. A wrong answer delivered confidently is worse than none.
Answered correctly, end to end
100%
ideal 100%
Judged on a sample of 16 by a small model, escalated to a larger one whenever the judge and the retrieval disagreed. Taken on Sep 6, 2026 — the judge is asked for, not automatic, so this is the last run that measured it rather than the last run.
Memory added per retrieval
1,709tokens
lower is better · has a floor
Roughly a thousand words of notes, carried into a conversation when it retrieves. The answer still has to fit, which is the floor.
Stayed quiet when it should have
12 of 12
Questions about things it has no notes on — sourdough, the World Cup, chess. Answering one from memory would be a false injection, and the other half of the guard.
Answered first try
88.89%
top three 95.56% of the time
The right note was the first thing returned. The strictest of the three bars, and the one with the most room left in it.
Questions in the test
45
The suite grew, which is why the lines above step down in the middle. The memory did not get worse; the test got harder.
The tiles above score a test. These score the automatic recall itself — what is handed to a working session before it asks anything, and what each question pulls in: a few past conversations and a few notes, each capped by length and re-ranked by a second, smaller model that reads the candidates once more and keeps the off-topic ones out.
Right conversation reached the session
90%
the best 3 · above a fixed floor · capped by length
The conversation half of automatic recall: the share of test questions where the right past conversation survived into what a session was actually handed.
Conversation memory added per question
574tokens
nine in ten questions get at most 581
The median, after the length cap. Add the notes half beside it for the real per-question bill.
Note memory added per question
522tokens
nine in ten get at most 601 · 2.6 notes each
The notes half, after re-ranking: up to 3 notes, each clipped to 900 characters. An off-topic question gets none.
Right note among the ones handed over
97.14%
the session's bar is 3, not five
Found-in-the-top-five is the test's bar; a session is handed three. This is recall at the bar that matters.
Read before any work starts
6,503tokens
the index alone is 3,711
Everything a session reads before its first question: the always-loaded index plus the two rule files. The standing tax the tuneups push down.
Costliest subject area to load
3,624tokens
the dearest of 12 areas
Loading an area means reading its routing block plus its three largest quick-reference blocks — the file path a session actually takes, which the retrieval scores never see. Which area it is stays unpublished.
Quick-reference blocks over budget
0
budget 0
Every note opens with a skim-only block capped at two kilobytes. One over the cap silently turns a cheap load into a full read.
Margin on the off-topic gate
+2.7pt
gate at -7 · worst off-topic -9.72
How far the best-scoring note on any off-topic question sat below the gate that keeps notes out. Positive means junk stays out; shrinking toward zero means the gate is drifting.
Clipped notes that lost the answer
0%
0 of 24 clipped right notes · held-out set 23.53% (4 of 17) · 66 of 91 notes ran past the 900-character cap
A note longer than the cap is cut around the bullet that best answers the question, not from the top. Of the right notes that needed cutting, this share still scored under the gate once cut — the answering line probably went with the tail. Cut from the top instead, the same run would have lost 20.83%.
Recall actually ran · last 7 days
100%
measured over 7 questions · handed memory on 100%
From the recall routine's own log. The rest saw a one-line note that memory was offline or busy that turn, and went on without it.
Held-out questions · right note in the top five
96.15%
95% band 81.11% – 99.32% · 26 questions
A second set nothing is ever tuned against, on Sep 6, 2026. Small, so the band matters more than the number; one run's move inside it is noise.
The first layer is one short file, read at the start of every session before anything else happens: the standing rules, what is open right now, and pointers into everything below. It is groomed by hand, and it regrows — the lines here are a month of that tug of war.
Size
14.5KB
1,858 words
The file itself, measured by a reporter that runs at the end of each session. The word count is the same file, counted at stamp time.
Rows
93lines
One line per pointer or rule. The count moves with the work; the size budget is what holds it.
Last compacted by hand
20days ago
on Aug 26, 2026
A tuneup, separate from grooming: closed work moved out, duplicates folded, the whole file re-read for what no longer earns its place.
Entries moved to the archive
53
28.4 KB · 143 lines
Finished work, kept in a companion file that loads only when a task touches it — so the always-loaded file carries only what is live.
The second layer keeps every past conversation, cut into pieces and embedded — turned into numbers that can be compared for meaning — in a Qdrant vector database, a store built for exactly that comparison. It fills itself, at the end of every session, with no one approving anything; that is why it is the high-recall layer and never the trusted one.
Searchable pieces of past conversation
4,748
+31 in the last week
Asked of the store directly. The line is the reporter's daily reading of the same count.
Working sessions
410
+3 in the last week
Distinct conversations in the archive, asked of it directly. The working copies on disk get cleaned up over time; the archive is the record that keeps everything.
Days of history
132days
earliest piece May 6, 2026 · latest Sep 15, 2026
From the date each piece carries in the store — the conversation's own date, never a file's last-touched time.
Storage segments
8
0 waiting to be ingested
How the store has split its index on disk. Furniture rather than a finding, published because the sibling dashboard shows it.
The third layer is the one worth trusting: one note per thing, holding how that thing actually works, plus written procedures for the tasks that recur. Every line in it was approved by a person, and each note carries the date it was last confirmed — so the layer knows its own age.
Confidence
- 98confirmed against the live thing
- 0inferred, awaiting confirmation
- 1stale — past the 90-day window
Notes, one per thing
99
Files in the notes folder, not counting the ones that only do routing. The line is the grooming routine's own count, one point per run.
Facts inside them
2,845
Every bullet across every note, counted by the grooming routine. A note is a file; a fact is a line in it.
Written procedures
25
Files in the procedures folder, counted the same way, with the month-end history asked of the repository.
Due for re-confirmation
1
past the 90-day window
Notes whose last confirmation is older than the window the house allows. The grooming routine lists them; a person re-checks them.
Subject areas the notes are grouped into
12
145 memberships · a note can sit in several
The table that defines them, so adding one updates this page by itself. How many, never which.
Size on disk
1.4MB
notes and procedures together
Every note and procedure file, measured by the grooming routine on its last run. The routing files and the private profile are left out.
Field reports published
3
One file per report in this site's own content folder
Commits
907
The repository's own history. Activity, never value — it is here because it carries the timeline on the Origin page, and for no other reason.
Routines it can be asked to run
40
Folders carrying a routine definition — a named task with its own instructions, invoked by name.
Machines it can log into
9
Counted from the notes, by what each note says it describes
Services it can deploy and reconfigure
39
The same count, on notes describing a service
Data stores it can read and write
11
The same count, on notes describing a store
Outside services it can reach
21
The same count, on notes describing an outside service
Stars on open-source projects
117,432,518stars
Across 15,434,222 projects, counted from the store directly. Only the one kind of activity this sensor collected in every month is counted; 23,524,832 rows of 13 other kinds, collected in a single wider sweep, are kept but left out — a month measured differently cannot be compared to the months beside it. Collection paused June 2024 – September 2025; the line skips the pause.
Collection stops here. The source withdrew the data.
Research papers
290,240papers
Counted from the instrument's own coverage record.
Company filings
19,572filings
Across 530 sources. Counted from the instrument's own coverage record.
Public posts
22,946posts
Across 2,359 sources. Counted from the instrument's own coverage record.
Recorded talks
3,148segments
Across 10 sources. Counted from the instrument's own coverage record.
What the market buys from AI models
11,145daily readings
One row per model per day, counted from the readings themselves rather than from the coverage record — this sensor holds readings of a moving market, not a pile of documents. Its window opens the day it was built; the market it reads keeps no archive, so there is no earlier history to collect.
The newest sensor reads the gateway this whole arrangement buys its own inference through, and records what the market did that day: which models the tokens went to, which companies served them, at what price, and how those models scored on third-party tests.
Every figure here except the first two is asked of the most recent sweep alone. These readings re-record the whole market daily, so counting across all of them would publish how many days it has run dressed up as the size of the market.
Models seen changing hands
513
Distinct models appearing in the daily readings between Aug 24, 2026 and Sep 13, 2026.
Days of readings held
21
Distinct days the market has been read. Low, and it can only ever grow forward — there is no history to buy.
Priced offers on the latest sweep
1,289
One model sold by one company at one price. Free and switched-off routes are excluded, because a route that costs nothing would flatter every price figure that follows.
Companies serving those models
106
The serving directory on the same sweep, counted directly.
Models carrying a capability score
96
Third-party test results, from the one source that publishes a comparable index. A second source ranks models against each other instead and is counted separately, never mixed in.
Trailing-window readings kept
35,688
Each sweep also snapshots the source's own rolling day, week, and month totals, over 21 sweeps. Stored against the day they were read, never the day they describe: a moving window read again tomorrow is a new measurement, not a correction.
Documents kept in full
1,322
The reading shelf, counted directly
Claims on the record
25
The claim ledger, counted directly. Claims that failed stay on it.
How those claims came out
14 mixed · 7 supported · 2 refuted · 1 untestable · 1 unverified
The same ledger, grouped by the verdict each claim was given. Refuted claims stay on it.
Model chosen to read filings
openai/gpt-oss-120b
5 candidate models read the same sample and were scored against the reference before this one was picked. It then read 18,260; the reference read 545.
Model chosen to read posts
deepseek/deepseek-v4-flash
The only candidate measured for this job — scored against the reference on a shared sample rather than picked from a field. It then read 9,570; the reference read 672.
Two rules about numbers
No figure that moves. Running costs, totals that only grow, anything that would have to be revisited to stay true — none of it is here. Every line on this page is redrawn by the script; none of them is edited.
No claims about how long anything took to build. They are hard to substantiate, they read as boasting, and speed is not what makes any of this worth trusting. There is a history on the Origin page with real dates on it; there is no stopwatch anywhere.
When the work happens
Everything is reported as a total, or at day precision. Time of day and day of the week say more about a household than about an instrument.
Which subjects the notes cover
How many, never which. The shape of what is known here is fair to publish; the index of it is not. The same rule hides which subject area is the costliest to load.
What any machine is called, or what it is for
Naming the technology is the point — it is what shows the work. Naming the machine is an inventory, and an inventory belongs on nobody's website.
What the outside review costs
The two reviewing models are metered by the token, and their bill is a running total. Their token counts are published; the dollars are not.