On a Tuesday in early spring, Jacob Coxon — until recently a safety researcher at Anthropic — published a short post on X. In it, he placed the probability that artificial intelligence eventually "kills all humanity" at something above ten percent. Hours earlier, a colleague at the same lab had resigned with a sentence I have not been able to put down: they are gambling with our lives. Coxon resigned too, and in his note he accused Anthropic and OpenAI of the same thing — racing toward a self-evolving superintelligence while the rest of us hold the collateral.
I have read many resignation letters in fourteen years of watching this industry. Most are administrative; they thank colleagues and mention new opportunities. This one was not, and what unsettled me was not the number. It was that nobody, including Coxon, can show us where the number comes from. A probability is not a proof; it is a confession. Confessions, however sincere, are the weakest class of evidence we accept.
I was twenty-one in 2017, a cryptography PhD candidate at UCL, auditing fifteen early-stage ICO whitepapers for a Medium series I called "The Soul of Code." Almost every one of them claimed to be secure, audited, and decentralized, and almost none of them showed a threat model. The tokenomics were structured so that speculation outran utility by construction, and the founders were sincere — that was the thing. Sincerity was not the failure. The absence of a mechanism was. From the chaos of 2017, we forged a compass, and the compass asked one question of every claim: show me the generative process, not the adjective.
The compass has served. During DeFi Summer I founded a Discord called The Trustless Circle for non-technical users, and we manually verified more than two hundred protocols against open-source standards, publishing a Trust Score dashboard; incident rates inside that community fell by roughly eighty percent not because the protocols improved but because people could finally read the risk. That is the whole lesson. Auditability is a design decision, and a lab that declines to be audited is not being modest. It is being unaccountable.
So when two safety researchers leave the most safety-branded laboratory in the world within the same week, I do not read the event as a warning about models. I read it as a warning about epistemics. And in the middle of an AI bull market — token narratives, compute contracts, valuations that quietly assume superintelligence arrives on schedule — the compass is pointing at something it cannot measure.
What does "greater than ten percent" actually mean? A probability claim is only as strong as the process that generates it: a base rate, a causal model, a simulation with published parameters, a red-team corpus with a defensible methodology. None of that exists in the public record for this number. There is no prior, no confidence interval, no falsifiable test that could prove Coxon wrong. A figure with no denominator is not a forecast; it is a moral gesture wearing the costume of arithmetic.
We solved a version of this problem in crypto, imperfectly but honestly. A rollup does not ask you to believe that a computation occurred; it hands you a fraud proof or a validity proof, and you check it yourself. The whole architecture is an admission that assertions are worthless without witnesses. Trust is not a metric; it is a memory we share — and the memory we are accumulating right now, across every accelerated training run, is that nobody outside the labs can verify anything at all.
This is why I built what I built. In 2026 I launched the Human-Centric AI Ledger, an initiative funded by ethical technology grants, to develop a cryptographic protocol that attests to the provenance of AI decisions: which model, which weights, which prompt chain, which human approvals — hashed, signed, and anchored so that an outside auditor can reconstruct the path after the fact. The goal is not to make intelligence safe by decree. The goal is to make the question "did this system do what it claimed?" answerable by somebody who does not work for the company that built it. That is all verification has ever meant.
Anthropic's constitutional AI and published model specifications are genuinely good artifacts — better than most of what the field produces. But a constitution that cannot be independently executed is a press release with footnotes. I have compared the published spec against behavior on adversarial prompts myself; you can see the seams where stated intent and measurable output diverge, and those seams are exactly where a resignation letter becomes data rather than testimony.
Now the harder thing. Coxon's ten percent is unauditable — and so is every reassurance either lab has ever offered. Both are narrative. We have no more public evidence that the race is safe than that it is lethal, and the asymmetry everyone feels is rhetorical, not empirical. I say this as someone who has spent a decade trying to convert trust into verification: emotional testimony and corporate safety blogs fail the same test, and they fail it in the same room. Two departures in one week is not an epistemology. It is a talent signal.
Which brings me to the blind spot. The deepest risk here is not that a model becomes superintelligent overnight; it is that safety becomes a brand asset — priced into a valuation, staffed generously while the narrative is hot, and quietly thinned the moment the competition accelerates. We watched this pathology in 2022, when projects collapsed not because their incentives were evil but because those incentives were misaligned with survival. Bull markets appreciate safety language much faster than they build safety systems. And when the people who do the actual adversarial work leave, the market does not reprice the risk. It reprices the story.
I keep thinking about blob space. Post-Dencun, everyone assumed cheap data availability was permanent; the data said it would saturate, and the fees came back. Prediction without verification is theater, and institutional memory is short.
So where does that leave a reader in the middle of a euphoric cycle? Not, I think, with a verdict on Anthropic or OpenAI, and certainly not with a comfortable one. It leaves us with a demand. Publish the methodology behind the ten percent, or retire the number. Publish the provenance of your models' decisions, or stop calling your deployments transparent. Verify the claim before you price it.
Ten percent is either a warning or a marketing position, and only one of those can survive an audit. The clock runs on the engineering, not the announcement.