The Findings
Claims you’ve already heard, taken seriously enough to measure.
Every report starts with something you’ve probably run into already. A pitch, a headline, someone fresh back from a conference. The claim sounds plausible and the stakes are real money, so instead of arguing with it, I measure it.
Same shape every time. The claim, then the public evidence read against it. Every finding leads with its takeaway, carries a grade, and ends with what I’d actually do about it. The report gets attacked before it goes up, and when a finding doesn’t survive that, it says so instead of quietly disappearing.
- Telling AI to think like Socrates: new voice, same answers, bigger billOn strong models, philosopher system prompts changed how the answers sounded and what they cost, not whether they were right. Every one of the 28 philosopher cases cost more than asking plainly, and no philosopher beat a same-length placebo prompt by more than the variation between two identical no-prompt runs. The one model that improved, a terse cheap one, improved under any long prompt.
- Buying an AI agent team: the manager isn't the expensive part“What you want is a manager agent directing worker agents across your tools and data, and the manager is the part worth paying up for.”
- Letting AI agents tune the system: adopted fast, measured almost never“Leave the decision engine alone and put a team of AI agents above it, offline, to tune it between runs.”
- What AI customer decisions should stand on“AI is going to make more of the decisions companies make about their customers, and you should get ahead of it.”