Skip to main content

Strength of Evidence

Every number in the corpus is first-party. That does not make them equally strong. Lead with the ones that disclose a method.

Hand-drawn line illustration in warm terracotta and cream: a balance scale where one solid cream block outweighs a pan of loose paper scraps, steadied by a hand and examined by an abstract one-line face.


Why tier at all

Quoting a directional productivity percentage to a skeptical engineer costs you the room. Quoting a benchmark with a disclosed method and a control group wins it. The difference is not the size of the number, it is whether the reader can tell how it was produced.

One useful fact about this corpus first: every figure comes from Anthropic's own customer-story pages, which are first-party company statements published with named attribution. No press coverage feeds it, so there is no journalist-versus-company reliability question to adjudicate. The distinctions below are entirely about method.

Tier A: controlled, benchmarked, or audited

Quote these first. Each discloses how the number was produced.

  • Smartsheet is the strongest structurally, because it is the only story in the corpus with a same-team control group: engineers using Claude Code compared against peers on the same teams. (source)
  • Freedom Forever published a real head-to-head with a stated methodology, building eight mock sites at five iterations each across 40 runs. (source)
  • eSentire evaluated against senior human analysts in production across more than 500 adjudicated outcomes. (source)
  • Rising Academies is the only externally validated result in the corpus, reporting an effect size from school-based studies run by researchers at Oxford and JPAL. (source)
  • Apollo ran blind testing and then a production A/B, attributing a retention change to the model swap alone. (source)
  • Semgrep, Graphite, Descript, Shortcut, Wordsmith, GC AI, and Satispay each disclose an evaluation set, a test-case count, or a structured comparison window.

Tier B: hard operational numbers

Self-reported, no disclosed method, but specific and falsifiable enough to be useful. These are counts and durations rather than percentages, which is what makes them checkable in principle.

Examples include Wiz on a line-count migration and the hours it took, Stripe on a language migration across a stated engineer count, Rakuten on time-to-market and an autonomous run length, LG CNS on APIs converted out of a total, Spotify and Delivery Hero on merged pull requests, Novo Nordisk on documentation turnaround, Dust on daily model spend, and Notion on cost and latency reduction from caching.

Tier C: credible directional numbers

The bulk of the corpus. Percentage productivity gains, time savings, adoption rates, none with a disclosed method. Useful for color and for showing a direction of travel. Always attribute them as what the company reports rather than as measured fact.

Tier D: use with care

Metrics the company itself flags as illustrative. At least one page prints a disclaimer stating its own quoted statistics are illustrative only and that results vary by configuration and context. Take the company at its word and do not quote those numbers as outcomes.

Projections rather than results. Several figures are forward-looking: revenue projected under an assumed monetization model, time savings a company expects rather than has recorded, hours it is on track to redirect. These are plans. Label them as plans.

Stories with no quantitative metrics at all. Roughly fourteen stories are qualitative only. They are still useful as existence proofs that a category of company shipped something, and useless as evidence of magnitude.

Pages that contradict themselves

Seven pages carry internal inconsistencies or cross-page contamination. Do not smooth these over. Pick one figure and cite the exact sentence it came from.

PageThe problem
BrexPrints two different expense-automation rates, and two different monthly hours-saved figures, on the same page.
ReplitHeadline and body disagree on user count, and the ARR sentence appears in two incompatible forms.
L'OréalAccuracy stated one way in the bullet list and another in the outcome section.
Qualified HealthHeadline callout and body bullet disagree on patient population size.
GraphiteThe top-of-page stat callouts belong to a different company's story entirely and match nothing in the body.
EmergentIts callout block carries an injected card containing another company's compliance metric.
ZoomA satisfaction figure appears on the page that belongs to a different customer's story.

The Graphite, Emergent, and Zoom cases are the same underlying defect: callout contamination across pages. If a headline statistic does not appear anywhere in the body prose, do not trust it.

One more attribution boundary

Several customer pages carry Anthropic marketing cards inside the content flow, making claims about enterprise AI adoption or time-to-production. Those are Anthropic's claims, not the customer's. When you quote a customer, quote the customer.

Further Reading