The Lab

The instrument the reports are readings from.

Ask a chatbot what people are saying about something and it will tell you, fluently, with no idea where the answer came from. This was built because that is not good enough to publish.


It is called Ouroboros, after the snake that eats its own tail, because what it settles gets written back into the thing it reads from. It has one job: to make a claim about what is happening in a field, and to make that claim traceable back to the rows it was counted from — so that a reader who doubts it can go and check rather than deciding whether to trust the tone.

The distinction it exists for is the difference between a doctor and someone playing a doctor on television. Both are fluent. Only one of them can show you why. This can still be wrong — and has been, on this page, below — but it is wrong in a way you can find.

Stars on open-source projects
Research papers
Company filings
Public posts
Recorded talks
What the market buys from AI models
SENSORSOne storeNOTHING DELETEDRead in stagesCHEAPEST FIRSTThe ledgerCLAIM FIRSTA field reportPUBLISHED HEREwhat was settled goes back in
Collection is the easy half. The two stages on the right are what separate a reading from an opinion: nothing reaches a report without a written claim behind it, and what the ledger settles is written back into the store, so the next question starts from a settled answer rather than from scratch.

What it watches

Six sensors, each collecting a different kind of public evidence, and each running on its own schedule into ClickHouse, a database built to count across hundreds of millions of rows in seconds rather than to serve an application. Nothing is ever deleted from it. That is a deliberate and slightly costly decision: you cannot go back and decide retroactively that something was worth keeping, so everything is kept.

SensorHoldsCovers
What people star in public117,432,518 starsacross 15,434,222 projectsJanuary 2023 – March 2026
What researchers publish290,240 papersJanuary 2023 – August 2026
What public companies tell their regulator19,572 filingsacross 530 companiesJanuary 2023 – August 2026
What people say about those projects22,946 postsacross 2,359 accountsSeptember 2025 – August 2026
What gets said out loud in recorded talks3,148 segmentsacross 10 channelsJanuary 2024 – August 2026
What the market buys from AI models11,145 daily readingsacross 513 modelsAugust 24, 2026 – September 13, 2026
Every window is published with the month it stops. The oldest sensor stopped collecting in March, when the source withdrew the data, and its window also holds a pause of more than a year in the middle. The newest one’s window opens the day it was built and is published in days rather than months, because the market it reads keeps no archive: nothing earlier exists to collect at any price. And the star count leaves out 23,524,832 rows the same sensor collected in January 2026 under a wider definition — 13 other kinds of activity, collected once and never again. They are still in the store. They are not in this figure, because a month measured differently cannot be compared to the months around it. All of it is a reason to distrust any claim the instrument makes about a stretch it did not collect.

Five of them watch what people say and make. The sixth watches what the market actually pays for — it reads the gateway the instrument itself buys its inference through, and records, every day, which models the world’s tokens went to, which companies served them, at what price, and how those models scored on third-party tests. It is the one sensor that cannot be caught up on later: the source serves a rolling window of about a month and sells no history at any price, so every day not collected is gone. Its readings are published as their own instrument, the model market.

Who does the reading

Most questions are answered by counting, which is what the database is for and costs nothing worth mentioning. The ones that are left need something to actually read a passage and judge what it says — and which model does that is a measured question here, not a preference.

Everything goes through one gateway — LiteLLM, which gives every provider the same shape so nothing downstream has to know which one it is talking to, pointed in turn at OpenRouter, which resells hundreds of models from one account. Between them, trying a different model is a line of configuration rather than a rewrite, and that is what makes the measurement practical: for each job, every candidate reads the same sample and is scored against a reference model before the winner does the bulk run. Some jobs draw a field of candidates; one had a single candidate, measured against the reference all the same rather than trusted on reputation. Different jobs end up with different winners, and the winner is rarely the expensive one.

Categorizing passages of company filings

Candidates measured
5
Chosen
openai/gpt-oss-120b
Read by the winner
18,260 passages
Read by the reference
545

Sorting what people said in public posts

Candidates measured
1
Chosen
deepseek/deepseek-v4-flash
Read by the winner
9,570 posts
Read by the reference
672
The reference model is the expensive one, and it never does the bulk run — it reads a sample twice to establish how much it agrees with itself, which is the ceiling every candidate is scored against, and then audits the winner. Paying up the price ladder has more than once bought nothing: on the filings, two models costing tens of times more were further from the reference than the one that won.

Alongside the sensors there is a shelf of 1,322 documents kept in full rather than counted — papers, posts, and articles worth reading properly rather than tallying. A number that comes out of a database query is only ever as good as the question; the shelf is what gets consulted when the question itself is the thing in doubt.

The claim gets written down first

Everything so far is collection and arithmetic, and neither of those makes anything trustworthy. What does is the ledger: the claim first, then the method for testing it, and only then is the test run. What the numbers say and what they are taken to mean are recorded as separate things, because the gap between those two is where most confident nonsense lives.

There are 25 claims on the record on the one subject it has been pointed at so far, of which 2 came back refuted and 14 came back mixed — the evidence pointing both ways at once. Claims that failed stay on the record. A ledger that quietly loses its wrong answers is a marketing document, and the whole point of writing the claim before the test is that it becomes impossible to pretend afterward that you expected the result you got.

What the first pass found

Companies have been putting real numbers behind their AI claims less and less over the last three years.

Clean, quotable, and consistent with what everyone already suspects. It is also an artifact of the paperwork: the mix of document sections shifted toward the ones that never quantify anything. Inside every individual section, the rate was flat.

What was actually happening

Nothing fell. The talk moved — out of the announcements companies make voluntarily and into the risks they are obliged to disclose.

A better finding than the one it replaced, and it only appeared because the first one had to be written down as a claim, tested inside each section separately, and allowed to fail.

Both statements are true about the same filings. Only the second one is about companies rather than about paperwork.

What gets published, and what does not

Most of what the ledger holds is not on this site, and will not be. A finding that cannot be stated without three qualifications is a finding that has not been understood yet. The reports carry the ones that survived, along with the limits of the data they came from — including, on more than one of them, the fields that were measured and then cut.

That is the honest summary of this whole section. There is a rented model in the middle of it, and around that model a set of notes, rules, and habits that between them can be pointed at a real question and produce an answer that does not have to be taken on faith. The reports are what that looks like when it works.