Provenance
Why you can trust this
This is not a chatbot's research summary. It is the output of a standing measurement instrument — and everything in it can be audited.
A fair first reaction to any AI-era research document is that someone typed a clever prompt into a model and formatted the answer. That is not what this is, and the difference is checkable. The numbers here are query results, not model opinions. Raw primary documents — SEC filings, software-contribution events, academic papers, price catalogs — were collected from their original sources into a permanent analytical database, and every figure attributed to our analysis is produced by versioned query code that can be re-run against that stored data, by anyone, for any audit.[1]
18,634SEC filings from 530 public companies, each traceable to its EDGAR accession number
29,287passages where those companies discuss AI under legal liability, 2023–present
141Mpublic software-development events behind the practitioner panel
66,704practitioners tracked month-by-month by what they contribute code to
290,240academic papers mapped to a shared concept vocabulary
406AI models price-snapshotted daily — a series that exists nowhere else
Where a language model was used, and where it was not. An LLM plays exactly one role in this pipeline: classifying filing passages into categories (does this passage state a dependency? quantify an outcome?). That step was treated as an instrument to be calibrated, not trusted: the classifier is an open-weight model pinned to a fixed prompt version, chosen by measuring candidate models against a frontier reference, and its noisiest term family was hand-audited on a 300-passage sample — 79% precision, published, and corrected for wherever it is used.[2] Every other number is deterministic SQL over stored data. No finding in this document is a model's summary of the internet.
Three disciplines stand behind every published number. Trends are re-run within each stratum, because a trend across a mixed population can be pure composition shift. Every "companies that do X also do Y" claim is re-tested within bands of how much each company writes, because verbose filers mention everything more. And on August 12, an adversarial re-analysis re-ran every load-bearing number from raw data with instructions to destroy it.[4] Three findings from the previous revision failed and were removed — they are named, with cause of death, in Removed & limits. What remains survived the attempt.
Provenance
How the instrument was tested
No sensor was trusted because it returned data. Each one was audited first — and the audits changed what could be claimed, killed one analysis outright, and refuted this project's own first headline.
Any measurement instrument is a claim about the world before it is a source of numbers, and the failure mode is specific: a broken sensor rarely returns nothing. It returns something plausible. So each source below was tested against something outside itself before any finding was allowed to rest on it, and the tests are listed here whether they passed or not.[19]
3independent confirmations required before accepting that a data source had been withdrawn upstream
0.96×a control that had to come out near 1.0, and did — the check that validates the arithmetic around it
1.03×what this project's own first headline measured out-of-sample, against 8.7× in-sample. Published as a refutation
18/18filing sections split correctly in a hand-audited pilot before the corpus was sectioned at scale
300passages hand-audited to measure the classifier's noisiest term family — 79% precision, corrected for
3findings killed by the adversarial re-analysis, each named in this report with its cause of death
- The upstream stream was audited before it was usedSignal-bearing event types in the public software-development stream fell about 85% between October 2025 and June 2026 — while the monthly total held flat near 110M, because push events grew 56% and absorbed the difference. Every day was present and the decay was smooth, so no coverage check trips on it, and a naive trend computed across that boundary would have reported that every project on earth was dying. Confirmed against the archive upstream of our own collection, establishing it as a real collection artifact rather than our loading error. The permanent consequence: the usable window ends around March 2026, and every score is a ranking within a period, never a level compared across periods.[19]
- Identity was tested, not assumedA project's name is not its identity. 124 of 1,840 repositories above 700 stars turned out to be aliases, and the rename rate rises with success — 10.3% above 10,000 stars against 5.9% in the 700–2,000 band. One project renamed twice in four days during a viral takeoff; keyed on its raw name it reads as three unremarkable mid-tier repositories instead of the largest single story in the window. Detecting renames structurally — one project's activity dying as another's is born — was tried first and rejected as far too noisy. The rule kept: structure may propose a candidate, the API decides.[20]
- A check that had to come out boring, and didThe crowd measure was run against the very project that defined the crowd — a number that must land near 1.0, because a crowd cannot discover what it has already discovered. It scored 0.96×. A pipeline with an error in it would have had no particular reason to land there. That the null case came out null is the strongest single reason to trust the arithmetic that surrounds it.[20]
- A source was withdrawn, and it took three independent checks to believe itThe original sensor was public star data. It stopped existing mid-2026 with no announcement, and "my pipeline broke" looks identical to "the source is gone" from inside a pipeline. Three checks, chosen to fail differently: the live event stream (star events fell from ~2M a month to ~20k while total events held steady), the historical archive sampled at the same hour across months (thousands per hour in March, roughly fifty by August, pushes flat), and the API endpoints themselves (star listings return not-found while the neighboring fork endpoint still answers). Only then was the instrument re-sensored.[21]
- The replacement sensor was scored against ground truth — and the first score was wrongValidated against the one month for which full firehose data was held, the new contribution sensor initially scored 74% recall. Investigating the apparent misses rather than accepting them: 62 of the 91 were projects held under a stale name. The sensor was not failing; the labels were. Resolved first, recall is 90%, and the new sensor surfaced 226 person-project pairs the old firehose never had. A validation number is itself a measurement, and it can be wrong in the direction that makes you abandon something that works.[12]
- The judge was graded before anything it judgedClassification uses a cheap open-weight model checked against a frontier reference. Before that, the reference was run twice on the same sample to see how often it agreed with itself: 92% on the load-bearing field — a real ceiling, now known — and 60% on another. That second field was not given a better prompt or a bigger model. It was dropped from the analysis, because a question the best available judge cannot answer consistently is not a measurement.[2]
- This project's own first headline was refuted by its own testAn earlier finding said a crowd of practitioners spotted breakout projects 8.7× more often than chance before takeoff. It survived every in-sample check. Then the method was frozen and pointed at a later window it had never seen, with a control matched on topic as well as activity — and the earliness it claimed to measure came out at 1.03×, indistinguishable from nothing. It was published as a refutation rather than quietly dropped. What survived is the part that replicated: the crowd is a 2.57× population filter, and it is not early.[22]
- An analysis was killed for being too thin to publishA planned cross-reference between the practitioner panel and social discussion found 13 of 2,990 post authors were panel members — 0.4%, a headcount rather than a trend, which more projects would not improve. The remaining collection work, roughly 35 hours, was cancelled rather than spent producing a weak signal that would have looked like a finding.[20]
- An audit pack was written for someone with no contextOn August 5 the whole instrument was documented for an outside second opinion: every figure queried live from the running system rather than recalled, a stated bias disclosure (it was written by the same agent that built the thing), and a section naming what its own author thought was weakest and most deserving of an outsider's skepticism.[20]
Provenance
How it was made, and what it took
The platforms, the models, and the effort — counted from the session logs, not estimated.
228human prompts across the 15 sessions that built and ran the instrument, Aug 3–13, 2026, counted from session transcripts
11days from an empty server to this report — database, four collectors, the analyses, and the writing
4revisions — three that removed or downgraded findings, one that changed only framing
1adversarial re-analysis: every load-bearing number re-derived from raw data with instructions to destroy it
$1.48total classification cost for the filing corpus, open-weight model
$8.93total spend on the instrument to date, against a self-imposed ceiling of $50 a month
The prompt count is the whole project, not the writing. It covers deciding which questions were worth asking, standing up the database, building and re-building the collectors, the validation passes listed above, the analyses that were killed, and the drafting — because quoting only the drafting sessions would describe a research document as though it were a blog post.[23]
Platforms. Collection and analysis run on a self-hosted stack: Python collectors feeding a ClickHouse analytical database, queried in SQL. Drafting, charting, and the adversarial re-analysis ran as Claude Code sessions against that database. The report itself is hand-authored HTML/MDX — no generator, no template engine.
Models, by role. Three models touched this report, in three different jobs. Claude Opus 5 and Claude Fable 5 (Anthropic) did the research direction, drafting, chart construction, and the red-team pass — under human instruction at every step. gpt-oss-120b (open-weight, run via API) did exactly one mechanical job: classifying 18,200 filing passages against a pinned prompt, selected for that job by measuring candidates against a frontier reference and costing $1.48 where the frontier model would have cost $212. No model generated any number cited as a finding — every figure is a SQL query result.
Provenance
Where AI was used — the full statement
Stated as method, not confessed as a caveat.
The prose of this report was written by an AI (Claude, Anthropic), working from query results, under human direction, across 22 prompts. The human — who owns the questions, the instrument, and every editorial judgment — reviewed each revision and ordered the adversarial pass that removed three findings the AI had drafted and initially defended.
The numbers were not produced by an AI. Every figure attributed to our analysis is the output of deterministic SQL over stored primary data. The single place a language model sits inside the measurement pipeline — passage classification — is treated as an instrument: pinned prompt, measured precision (79% on its noisiest family, hand-audited on 300 passages), error corrected for where it is used, and any field the reference model could not answer consistently (below 92% self-agreement) dropped entirely.
What that means for the reader: you are trusting the stored data, the published query code, and the stated calibration — not a model's recollection of the internet. All three can be audited; the report's citations say where.
Provenance
Version history
This page renders Revision 4, unabridged. Earlier revisions are superseded, not hidden — what each one lost is listed in Removed & limits.
- Rev 1 · Aug 11, 2026First full draft: 17 candidate findings from the four evidence bases. Two findings removed the same day under standard controls (verbosity, composition).
- Rev 2 · Aug 12, 2026The adversarial re-analysis — every load-bearing number re-derived from raw data by a fresh session instructed to destroy the report. Three findings failed and were removed; two more were re-scoped. The surviving content is what this page shows.
- Rev 3 · Aug 12, 2026Content unchanged from Revision 2 except where noted in the text. Adds navigation, per-claim citations, charts, and the provenance summary.
- Rev 4 · Aug 13, 2026A framing pass only. Earlier revisions addressed a single organization in the first person; this one addresses any reader facing the same question. No finding, number, grade, confounder or citation changed — the evidence is identical to Revision 3.
Every revision has removed or downgraded something. None added reach. What each one lost is named, with cause of death, in Removed & limits.