The Lab

Every number, and where it came from.

Claims about an AI setup are cheap, so this tab only carries figures that can be counted from something. Every reading below says what it was counted from, and every line that moves is drawn from a real history rather than a guess.


All of it was counted on September 15, 2026 by a script in the workshop repository, run every time the notes are groomed, a session is closed out, or the memory is tuned. Nothing on any of these pages is typed in by hand — and where a source has no honest history, there is no line drawn, rather than a line invented.

The first five panels mirror the dashboard the agent’s own control panel shows its author, panel for panel and in the same order. What that dashboard shows and this page does not: the name of any machine, the name of any subject, and any bill.

How the memory gets written36 runs since Jun 10, 2026aging

Last groomed

9days ago

September 6, 2026

Written only by the grooming routine, only after a person approves

9revised1reconfirmed2private-profile edits1procedure touched

Nothing writes into long-term memory on its own. A grooming routine reads the sessions since it last ran, works out what is worth keeping, and produces a proposal — notes to add, notes whose facts have changed, notes it now doubts. Then it stops. The proposal is applied only when a person says so, and every run is recorded: the memory has a change log, and the counts here are that log rather than an intention.

Grooming runs

36

Between Jun 10, 2026 and Sep 6, 2026. One row per proposal.

Notes revised

120

A fact that turned out to be wrong or out of date, corrected. Far more common than writing a new note.

Notes added

7

New things worth a note of their own. The low number is the point: most learning corrects something rather than adding to it.

Procedures touched

12

Written procedures added or revised by a grooming run, over the same window.

Outside review6 runs reviewed · 148 items · 0 corrections accepted

Two other models — Gemini and GPT, rented the same way the main one is — read the same proposal before it is applied. On each item a reviewer either agrees with the groomer or files a disagreement; a disagreement whose fix is better goes into memory, and one that is not is overruled. A run where both simply agreed is the healthy default, not a failing score.

Gemini 2.5 Pro

Agreed with the groomer on all 13 items

  • 12agreed with the groomer
  • 0disagreed · the groomer's version kept
  • 0disagreed · the reviewer's fix accepted

of 13 reviewed21.3K tokens read · 3.7K written

GPT-4.1

Agreed with the groomer on all 13 items

  • 13agreed with the groomer
  • 0disagreed · the groomer's version kept
  • 0disagreed · the reviewer's fix accepted

of 13 reviewed18.2K tokens read · 1.9K written

How the memory is tested14 runs · latest Sep 6, 2026guards holding

The memory is checked against a set of questions whose right answer is already known — questions this house has genuinely had to answer, each one tied to the note that should come back. A run asks all of them and scores what came back against what should have.

Mixed in are 12 questions about subjects it holds no notes on at all. Those are the important ones. A memory that answers everything confidently has stopped being a memory, so staying quiet on those is scored as getting them right.

Every one of the 124 notes and procedures can be reached by at least one question in the test. A note nothing can reach cannot miss, so this line comes before any score.

Right note in the top five

100%

ideal 100%

Per run · 14 runs

Across 45 questions with known answers, on Sep 6, 2026.

Ranking quality

93.22%

a score, not a share · first 88.89% of the time

Per run · 14 runs

Highest when the right note is first and falling away as it slips down the list. Read it against the two tiles beside it, never as a percentage of anything.

Of the five returned, share that were right

41.15%

low by design · about one right answer per question

Most questions have one right note, so five slots holding one correct answer keeps this near twenty percent — it runs higher only where a question has more than one. It is here because leaving it out would look like hiding it.

Served an out-of-date note

0%

must stay 0%

A note that has since been superseded, handed over as if current. A wrong answer delivered confidently is worse than none.

Answered correctly, end to end

100%

ideal 100%

Judged on a sample of 16 by a small model, escalated to a larger one whenever the judge and the retrieval disagreed. Taken on Sep 6, 2026 — the judge is asked for, not automatic, so this is the last run that measured it rather than the last run.

Memory added per retrieval

1,709tokens

lower is better · has a floor

Roughly a thousand words of notes, carried into a conversation when it retrieves. The answer still has to fit, which is the floor.

Stayed quiet when it should have

12 of 12

Questions about things it has no notes on — sourdough, the World Cup, chess. Answering one from memory would be a false injection, and the other half of the guard.

Answered first try

88.89%

top three 95.56% of the time

Per run · 14 runs

The right note was the first thing returned. The strictest of the three bars, and the one with the most room left in it.

Questions in the test

45

Per run · 14 runs

The suite grew, which is why the lines above step down in the middle. The memory did not get worse; the test got harder.

What a session actually getsMeasured on the path a session takes, not the test's

The tiles above score a test. These score the automatic recall itself — what is handed to a working session before it asks anything, and what each question pulls in: a few past conversations and a few notes, each capped by length and re-ranked by a second, smaller model that reads the candidates once more and keeps the off-topic ones out.

Right conversation reached the session

90%

the best 3 · above a fixed floor · capped by length

The conversation half of automatic recall: the share of test questions where the right past conversation survived into what a session was actually handed.

Conversation memory added per question

574tokens

nine in ten questions get at most 581

The median, after the length cap. Add the notes half beside it for the real per-question bill.

Note memory added per question

522tokens

nine in ten get at most 601 · 2.6 notes each

The notes half, after re-ranking: up to 3 notes, each clipped to 900 characters. An off-topic question gets none.

Right note among the ones handed over

97.14%

the session's bar is 3, not five

Found-in-the-top-five is the test's bar; a session is handed three. This is recall at the bar that matters.

Read before any work starts

6,503tokens

the index alone is 3,711

Everything a session reads before its first question: the always-loaded index plus the two rule files. The standing tax the tuneups push down.

Costliest subject area to load

3,624tokens

the dearest of 12 areas

Loading an area means reading its routing block plus its three largest quick-reference blocks — the file path a session actually takes, which the retrieval scores never see. Which area it is stays unpublished.

Quick-reference blocks over budget

0

budget 0

Every note opens with a skim-only block capped at two kilobytes. One over the cap silently turns a cheap load into a full read.

Margin on the off-topic gate

+2.7pt

gate at -7 · worst off-topic -9.72

How far the best-scoring note on any off-topic question sat below the gate that keeps notes out. Positive means junk stays out; shrinking toward zero means the gate is drifting.

Clipped notes that lost the answer

0%

0 of 24 clipped right notes · held-out set 23.53% (4 of 17) · 66 of 91 notes ran past the 900-character cap

A note longer than the cap is cut around the bullet that best answers the question, not from the top. Of the right notes that needed cutting, this share still scored under the gate once cut — the answering line probably went with the tail. Cut from the top instead, the same run would have lost 20.83%.

Recall actually ran · last 7 days

100%

measured over 7 questions · handed memory on 100%

From the recall routine's own log. The rest saw a one-line note that memory was offline or busy that turn, and went on without it.

Held-out questions · right note in the top five

96.15%

95% band 81.11% – 99.32% · 26 questions

A second set nothing is ever tuned against, on Sep 6, 2026. Small, so the band matters more than the number; one run's move inside it is noise.

Retrieval quality over runseach on its own scale · change across the last 14 runs
In the top five+6.7 pt · 14 runs
100%ideal 100%
In the top three+5.6 pt · 14 runs
95.56%ideal 100%
First returned+28.9 pt · 14 runs
88.89%higher is sharper
Ranking quality+19.1 pt · 14 runs
93.22%100% would be always first
Memory · the always-loaded indexAs of Sep 15, 2026aging

The first layer is one short file, read at the start of every session before anything else happens: the standing rules, what is open right now, and pointers into everything below. It is groomed by hand, and it regrows — the lines here are a month of that tug of war.

Size

14.5KB

1,858 words

Per day · Aug 16, 2026 – Sep 15, 2026

The file itself, measured by a reporter that runs at the end of each session. The word count is the same file, counted at stamp time.

Rows

93lines

Per day · Aug 16, 2026 – Sep 15, 2026

One line per pointer or rule. The count moves with the work; the size budget is what holds it.

Last compacted by hand

20days ago

on Aug 26, 2026

A tuneup, separate from grooming: closed work moved out, duplicates folded, the whole file re-read for what no longer earns its place.

Entries moved to the archive

53

28.4 KB · 143 lines

Finished work, kept in a companion file that loads only when a task touches it — so the always-loaded file carries only what is live.

30pinned rules9open work0gotchas kept here26reference pointers53archived
Memory · the searchable archiveAs of Sep 15, 2026healthy

The second layer keeps every past conversation, cut into pieces and embedded — turned into numbers that can be compared for meaning — in a Qdrant vector database, a store built for exactly that comparison. It fills itself, at the end of every session, with no one approving anything; that is why it is the high-recall layer and never the trusted one.

Searchable pieces of past conversation

4,748

+31 in the last week

Per day · Aug 16, 2026 – Sep 15, 2026

Asked of the store directly. The line is the reporter's daily reading of the same count.

Working sessions

410

+3 in the last week

Per day · Aug 16, 2026 – Sep 15, 2026

Distinct conversations in the archive, asked of it directly. The working copies on disk get cleaned up over time; the archive is the record that keeps everything.

Days of history

132days

earliest piece May 6, 2026 · latest Sep 15, 2026

From the date each piece carries in the store — the conversation's own date, never a file's last-touched time.

Storage segments

8

0 waiting to be ingested

How the store has split its index on disk. Furniture rather than a finding, published because the sibling dashboard shows it.

Memory · the notesAs of Sep 15, 2026reconfirm soon

The third layer is the one worth trusting: one note per thing, holding how that thing actually works, plus written procedures for the tasks that recur. Every line in it was approved by a person, and each note carries the date it was last confirmed — so the layer knows its own age.

Confidence

  • 98confirmed against the live thing
  • 0inferred, awaiting confirmation
  • 1stale — past the 90-day window

Notes, one per thing

99

Per reading · last 50

Files in the notes folder, not counting the ones that only do routing. The line is the grooming routine's own count, one point per run.

Facts inside them

2,845

Per reading · last 50

Every bullet across every note, counted by the grooming routine. A note is a file; a fact is a line in it.

Written procedures

25

Month end · June 2026 – September 2026

Files in the procedures folder, counted the same way, with the month-end history asked of the repository.

Due for re-confirmation

1

past the 90-day window

Notes whose last confirmation is older than the window the house allows. The grooming routine lists them; a person re-checks them.

Subject areas the notes are grouped into

12

145 memberships · a note can sit in several

The table that defines them, so adding one updates this page by itself. How many, never which.

Size on disk

1.4MB

notes and procedures together

Every note and procedure file, measured by the grooming routine on its last run. The routing files and the private profile are left out.

The workshopSince April 2026

Field reports published

3

One file per report in this site's own content folder

Commits

907

Per month · April 2026 – September 2026

The repository's own history. Activity, never value — it is here because it carries the timeline on the Origin page, and for no other reason.

Routines it can be asked to run

40

Folders carrying a routine definition — a named task with its own instructions, invoked by name.

ReachCounts, never names

Machines it can log into

9

Counted from the notes, by what each note says it describes

Services it can deploy and reconfigure

39

The same count, on notes describing a service

Data stores it can read and write

11

The same count, on notes describing a store

Outside services it can reach

21

The same count, on notes describing an outside service

The instrument · sensorsEach line spans only the months that sensor actually collected

Stars on open-source projects

117,432,518stars

January 2023 – March 2026

Across 15,434,222 projects, counted from the store directly. Only the one kind of activity this sensor collected in every month is counted; 23,524,832 rows of 13 other kinds, collected in a single wider sweep, are kept but left out — a month measured differently cannot be compared to the months beside it. Collection paused June 2024 – September 2025; the line skips the pause.

Collection stops here. The source withdrew the data.

Research papers

290,240papers

January 2023 – August 2026

Counted from the instrument's own coverage record.

Company filings

19,572filings

January 2023 – August 2026

Across 530 sources. Counted from the instrument's own coverage record.

Public posts

22,946posts

September 2025 – August 2026

Across 2,359 sources. Counted from the instrument's own coverage record.

Recorded talks

3,148segments

January 2024 – August 2026

Across 10 sources. Counted from the instrument's own coverage record.

What the market buys from AI models

11,145daily readings

August 24, 2026 – September 13, 2026

One row per model per day, counted from the readings themselves rather than from the coverage record — this sensor holds readings of a moving market, not a pile of documents. Its window opens the day it was built; the market it reads keeps no archive, so there is no earlier history to collect.

The instrument · the inference marketLatest sweep · Sep 13, 2026

The newest sensor reads the gateway this whole arrangement buys its own inference through, and records what the market did that day: which models the tokens went to, which companies served them, at what price, and how those models scored on third-party tests.

Every figure here except the first two is asked of the most recent sweep alone. These readings re-record the whole market daily, so counting across all of them would publish how many days it has run dressed up as the size of the market.

Models seen changing hands

513

Distinct models appearing in the daily readings between Aug 24, 2026 and Sep 13, 2026.

Days of readings held

21

Distinct days the market has been read. Low, and it can only ever grow forward — there is no history to buy.

Priced offers on the latest sweep

1,289

One model sold by one company at one price. Free and switched-off routes are excluded, because a route that costs nothing would flatter every price figure that follows.

Companies serving those models

106

The serving directory on the same sweep, counted directly.

Models carrying a capability score

96

Third-party test results, from the one source that publishes a comparable index. A second source ranks models against each other instead and is counted separately, never mixed in.

Trailing-window readings kept

35,688

Each sweep also snapshots the source's own rolling day, week, and month totals, over 21 sweeps. Stored against the day they were read, never the day they describe: a moving window read again tomorrow is a new measurement, not a correction.

The instrument · reading and judgingNothing here is a running total

Documents kept in full

1,322

The reading shelf, counted directly

Claims on the record

25

The claim ledger, counted directly. Claims that failed stay on it.

How those claims came out

14 mixed · 7 supported · 2 refuted · 1 untestable · 1 unverified

The same ledger, grouped by the verdict each claim was given. Refuted claims stay on it.

Model chosen to read filings

openai/gpt-oss-120b

5 candidate models read the same sample and were scored against the reference before this one was picked. It then read 18,260; the reference read 545.

Model chosen to read posts

deepseek/deepseek-v4-flash

The only candidate measured for this job — scored against the reference on a shared sample rather than picked from a field. It then read 9,570; the reference read 672.

Two rules about numbers

No figure that moves. Running costs, totals that only grow, anything that would have to be revisited to stay true — none of it is here. Every line on this page is redrawn by the script; none of them is edited.

No claims about how long anything took to build. They are hard to substantiate, they read as boasting, and speed is not what makes any of this worth trusting. There is a history on the Origin page with real dates on it; there is no stopwatch anywhere.

What is deliberately not countedFour more, beyond the two rules above

When the work happens

Everything is reported as a total, or at day precision. Time of day and day of the week say more about a household than about an instrument.

Which subjects the notes cover

How many, never which. The shape of what is known here is fair to publish; the index of it is not. The same rule hides which subject area is the costliest to load.

What any machine is called, or what it is for

Naming the technology is the point — it is what shows the work. Naming the machine is an inventory, and an inventory belongs on nobody's website.

What the outside review costs

The two reviewing models are metered by the token, and their bill is a running total. Their token counts are published; the dollars are not.