The instrument the reports are readings from.
Ask a chatbot what people are saying about something and it will tell you, fluently, with no idea where the answer came from. This was built because that is not good enough to publish.
It is called Ouroboros, after the snake that eats its own tail, because what it settles gets written back into the thing it reads from. It has one job: to make a claim about what is happening in a field, and to make that claim traceable back to the rows it was counted from — so that a reader who doubts it can go and check rather than deciding whether to trust the tone.
The distinction it exists for is the difference between a doctor and someone playing a doctor on television. Both are fluent. Only one of them can show you why. This can still be wrong — and has been, on this page, below — but it is wrong in a way you can find.
What it watches
Five sensors, each collecting a different kind of public evidence, and each running on its own schedule into ClickHouse, a database built to count across hundreds of millions of rows in seconds rather than to serve an application. Nothing is ever deleted from it. That is a deliberate and slightly costly decision: you cannot go back and decide retroactively that something was worth keeping, so everything is kept.
| Sensor | Holds | Covers |
|---|---|---|
| What people build in public | 140,957,350 activity recordsacross 17,087,365 projects | January 2023 – March 2026 |
| What researchers publish | 290,240 papers | January 2023 – August 2026 |
| What public companies tell their regulator | 19,572 filingsacross 530 companies | January 2023 – August 2026 |
| What people say about those projects | 22,946 postsacross 2,359 accounts | September 2025 – August 2026 |
| What gets said out loud in recorded talks | 3,148 segmentsacross 10 channels | January 2024 – August 2026 |
Who does the reading
Most questions are answered by counting, which is what the database is for and costs nothing worth mentioning. The ones that are left need something to actually read a passage and judge what it says — and which model does that is a measured question here, not a preference.
Everything goes through one gateway — LiteLLM, which gives every provider the same shape so nothing downstream has to know which one it is talking to, pointed in turn at OpenRouter, which resells hundreds of models from one account. Between them, trying a different model is a line of configuration rather than a rewrite, and that is what makes the measurement practical: for each job, every candidate reads the same sample and is scored against a reference model before the winner does the bulk run. Some jobs draw a field of candidates; one had a single candidate, measured against the reference all the same rather than trusted on reputation. Different jobs end up with different winners, and the winner is rarely the expensive one.
Categorizing passages of company filings
- Candidates measured
- 5
- Chosen
- openai/gpt-oss-120b
- Read by the winner
- 18,260 passages
- Read by the reference
- 545
Sorting what people said in public posts
- Candidates measured
- 1
- Chosen
- deepseek/deepseek-v4-flash
- Read by the winner
- 9,570 posts
- Read by the reference
- 672
Alongside the sensors there is a shelf of 65 documents kept in full rather than counted — papers, posts, and articles worth reading properly rather than tallying. A number that comes out of a database query is only ever as good as the question; the shelf is what gets consulted when the question itself is the thing in doubt.
The claim gets written down first
Everything so far is collection and arithmetic, and neither of those makes anything trustworthy. What does is the ledger: the claim first, then the method for testing it, and only then is the test run. What the numbers say and what they are taken to mean are recorded as separate things, because the gap between those two is where most confident nonsense lives.
There are 22 claims on the record on the one subject it has been pointed at so far, of which 1 came back refuted and 14 came back mixed — the evidence pointing both ways at once. Claims that failed stay on the record. A ledger that quietly loses its wrong answers is a marketing document, and the whole point of writing the claim before the test is that it becomes impossible to pretend afterward that you expected the result you got.
What the first pass found
Companies have been putting real numbers behind their AI claims less and less over the last three years.
Clean, quotable, and consistent with what everyone already suspects. It is also an artifact of the paperwork: the mix of document sections shifted toward the ones that never quantify anything. Inside every individual section, the rate was flat.
What was actually happening
Nothing fell. The talk moved — out of the announcements companies make voluntarily and into the risks they are obliged to disclose.
A better finding than the one it replaced, and it only appeared because the first one had to be written down as a claim, tested inside each section separately, and allowed to fail.
What gets published, and what does not
Most of what the ledger holds is not on this site, and will not be. A finding that cannot be stated without three qualifications is a finding that has not been understood yet. The reports carry the ones that survived, along with the limits of the data they came from — including, on more than one of them, the fields that were measured and then cut.
That is the honest summary of this whole section. There is a rented model in the middle of it, and around that model a set of notes, rules, and habits that between them can be pointed at a real question and produce an answer that does not have to be taken on faith. The reports are what that looks like when it works.