The Sequel Nobody Writes
Every year the AI-and-jobs discourse produces a fresh crop of numbers with dates attached. Half of entry-level white-collar work, gone in five years. Ninety-two million jobs displaced by 2030. Forty-seven percent of American employment at risk. Almost none of them are ever checked.
That is not an accident of laziness. It is structural: the institutions that publish forecasts are the same ones that publish the next forecast, and a retrospective is a bad look with no audience. When the World Economic Forum’s 2020 projection — 85 million jobs displaced by 2025 — reached its deadline, the WEF’s response was to publish a new pair of numbers for 2030 without grading the old pair. Nobody else graded them either. We looked.
This is a scoreboard. Fifty-five dated claims about AI and work, each one graded against evidence published since our own piece on the subject went up on 18 July 2026 — on track, overstated, understated, or too early to call, with the reasoning and the source behind every verdict.
Ten of the rows are ours.
Verdicts are our reading of the evidence, not a measurement. Colour is a reading aid — every verdict is also named in text, and the table view carries the whole board without colour.
Graded 7 September 2026 against evidence published since 18 July 2026.
What Fifty-Five Verdicts Look Like
The single clearest pattern is that the failure mode is timing, not direction.
Twenty-three claims are on track and seventeen are overstated — but read the overstated ones and almost every case is a forecaster who described the right shape and the wrong year. Geoffrey Hinton said in 2016 that deep learning would beat radiologists within five years and that people should stop training them. Radiologist headcount, residency places, salaries and unfilled vacancies have all risen since; there are 4,333 open posts taking an average of 130 days to fill. But AI is now genuinely good at reading medical images, and Hinton’s own concession is the honest summary: “wrong on the timing but not the direction.”
The same shape recurs. Elon Musk promised a million robotaxis on the road in 2020; there were none, and as of September 2026 Tesla has 420 vehicles registered for automated operation in Texas. Yet Waymo runs paid driverless service in fourteen cities at over half a million rides a week. The technology milestone largely arrived. The employment consequence has not: US truck driver employment rose about 30% over the decade, and the government’s own occupational outlook for the job doesn’t mention autonomous vehicles at all.
Capability and labour-market impact came apart by more than a decade. Every forecast that assumed they move together got the year wrong, and most of them got it wrong in the same direction.
Three Things Worth Knowing Before You Cite Anything
A lot of the famous numbers were never forecasts. Frey and Osborne’s “47% of US jobs” is the most-cited statistic in this entire field, and page 44 of their paper says: “We make no attempt to estimate the number of jobs that will actually be automated.” It measures technical automatability, with no adoption pace and no employment prediction. Frey has spent a decade saying so. The Obama White House’s 2.2-to-3.1-million driving-jobs figure carries an explicit note that it “is not a net calculation”. Goldman Sachs’s 2017 driver-displacement note put the full effect “several decades away”. All three are routinely quoted as near-term job-loss predictions. None of them were.
Nobody official attributes job losses to AI. When the US Bureau of Labor Statistics finally built an AI-exposure measure in August 2026 — using, among other things, Anthropic’s and Microsoft’s own usage data — it wrote the disclaimer into the documentation: “An exposure category is not a forecast of employment growth or decline. Exposure does not imply job loss.” Every AI-attribution number in circulation is a projection, an exposure score, or an employer’s self-reported reason for a layoff. The last of those is the weakest and the most quoted.
Check who is grading whom. The most creditworthy evidence in this dataset is the evidence that embarrassed its publisher. Microsoft’s 2025 Work Trend Index found 81% of leaders expecting AI agents to be moderately or extensively integrated within 12–18 months; Microsoft’s own 2026 edition, fielded as that window closed, is built around what it calls the Transformation Paradox — “Workers are ready. Their organizations aren’t.” Only 19% of AI users sit in its Frontier zone. Even in software and technology, the leading industry, roughly one firm in five uses agents at all. Meanwhile Anthropic’s head of economics spent July 2026 publicly contradicting Anthropic’s chief executive: “I don’t expect unemployment to be noticeably higher a year from now — at least not because of AI.” US unemployment was 4.1% in August, against a forecast of 10–20%.
The Entry-Level Question, Carefully
The claim we made loudest last year was that it’s the entry level that’s coughing, not the whole mine. That holds — and the evidence has become considerably more interesting than a simple yes.
Stanford’s Canaries in the Coal Mine is the most-cited empirical result in this debate, and its August 2026 revision reports the gap between young workers in AI-exposed occupations and their less-exposed peers widening to 19%. But the headline number is not one series: the earlier 13% and 16% figures were firm-shock-adjusted regression estimates, while the 19% is an uncontrolled descriptive divergence. On a consistent basis the paper back-casts 15% rising to 19%. Buried in its own appendix is the disclosure that on the original filters, the within-firm estimate for the most-exposed quintile fell from −11.7 log points to −5.3 and lost statistical significance. It is still a working paper, still not peer-reviewed, on proprietary data the authors cannot share.
Meanwhile Handshake — the broadest new-graduate dataset there is — shows postings down 2% year over year. US unemployment for 20-to-24-year-olds has fallen from 9.2% to 7.1% since September 2025. The OECD’s July 2026 Employment Outlook concludes that large language models are not the main cause of graduate unemployment at all, because the gap has been widening since before the pandemic with no turning point at the LLM inflection.
All of these are true at once, and they are not in contradiction. Stanford measures relative employment between exposed and unexposed occupations, not levels. Something real and specific is happening at the bottom of specific ladders. It is not an economy-wide entry-level collapse, and the shape that is actually emerging is stranger than either: PwC finds AI-exposed entry roles are seven times more likely to demand senior-level skills, with “seniorised” entry jobs up 35% since 2019 while other entry-level roles fell 10%. Indeed calls the pattern seniority-biased technological change. The bottom rung isn’t disappearing. It’s being raised out of reach.
Grading Ourselves
Self-scoring is what separates a scoreboard from a scolding, so ten rows on the board are ours, graded on the same axis as everyone else’s. Three hold up, five are overstated — two of them outright factual errors — and two are too early to call.
We wrote that the OECD finds “the largest, most stubborn gaps in manipulation (0.7) and robotic intelligence (0.6)”. The OECD’s largest gaps are in social interaction, problem solving and metacognition, all at 1.1; the physical domains rank third and fourth. What the report actually says about them is that they “may prove the most stubborn to close” — a claim about difficulty, which we turned into a claim about size. We also attributed the index’s highest score, 6.4, to “personal care, social work and community services”; it belongs to community and social service occupations alone, and personal care sits on the exposed side. Both are now corrected in the source piece.
The more interesting failure is the argument rather than the arithmetic. We called augmentation-versus-automation “the closest thing to a unifying theory” in the research. Stanford’s own August 2026 data shows the complementarity coefficient for 22-to-25-year-olds is insignificant and slightly negative — the protective effect of augmentation appears only for workers over 40, which is to say it does not hold for the group the whole piece was worried about. And we quoted Anthropic’s consumer figure, 52% augmented against 45% automated, without noticing that Anthropic’s enterprise API traffic — where firms actually deploy AI against payroll — runs about 75% automated, and that Anthropic explicitly reads work migrating there as a leading indicator of displacement. We picked the friendlier half of the data. The distinction is real and still useful. Calling it a unifying theory was a reach.
The thing that graded best was the section where we admitted what we didn’t know. All three of its hedges held: the Stanford paper we flagged as un-peer-reviewed still is, the Microsoft caveat we repeated has since been vindicated by three independent datasets, and the general warning that these are all bets on trend lines turns out to be the finding of this entire exercise.
Built to Be Re-Run
This page is designed to be graded again next September, and the year after. The claims are historical facts and never change; what changes is a verdict, a paragraph, and a citation. The dataset lives apart from the prose specifically so that the second edition is an afternoon’s work rather than a rewrite — and so that when a verdict flips, the flip is legible.
Two rows are worth watching in particular. Amodei’s 10–20% unemployment window runs to May 2030, and is currently sixteen points out and moving away. Metaculus forecasters expect US employment to fall about 3% over the next decade against a government projection of +3% — a six-point disagreement with a public, dated resolution method, which is the one thing essentially none of the forecasts graded here ever had.
If there’s a single lesson in fifty-five verdicts, it’s the one our own humility section stumbled into: hold a forecast the way you’d hold any honest bet, and pay more attention to who states their mechanism than to who states the biggest number. The forecasts that graded well — Arntz’s 9%, the IMF’s complementarity axis, Microsoft’s own caveat against its own data — all did the same unglamorous thing. They said what would have to be true, and how you’d know.