CoachDiscoverMapCompareSavedAI LearningAI WeeklyAbout
← Back to library
01 · Foundations

What the Stanford Study Found

51 enterprise AI deployments that cleared a strict bar — live, consistent, measurable, scalable. The 5% reverse-engineered from a 95% pilot graveyard. One thesis: the model wasn't the moat.

The strict bar
Live in production · Used consistently · Measurable business value · Capable of scaling. All four — or it's not in the dataset.
The 95/5 reality
Roughly 95% of GenAI pilots fail to produce measurable financial impact. The 51 are the reverse-engineered profile of the 5%.
The thesis
"The difference was never the AI model. It was always the organization." Readiness, processes, leadership, willingness to redesign and absorb failure.

The bar — what counts as "successful"

Most enterprise AI surveys count pilots that "showed promise" or "generated learnings." Stanford's playbook used a much stricter inclusion test, which is why its 51 cases describe a smaller and more useful population.

Criterion
Why it matters
What gets excluded
Live in production
Pilot ≠ product. A demo that never shipped is not a deployment.
Side-pocket experiments, sandbox demos.
Used consistently
An end-of-quarter usage spike doesn't prove adoption.
Tools nobody opens after week 4.
Measurable business value
"Felt productive" is not a unit of output.
Cases without baseline + delta numbers.
Capable of scaling further
A one-team win that won't generalize is a curiosity.
Bespoke wins with no scaling path.
The four-part bar is what gives the rest of the playbook its weight. The findings aren't about "what enterprise AI looks like." They're about what survived to production and stayed there.

The 51 in numbers

A five-month empirical study by Stanford's Digital Economy Lab. The dataset is wide enough to compare across sectors, deep enough to compare across phases of the same deployment lifecycle.

Dimension
Count
Why it matters
Deployments
51
Enough to see patterns across cases, not enough to hide in averages.
Organizations
41
Some orgs had multiple deployments — cross-deployment learnings within one org.
Industries
9
Findings hold across sector, not just tech.
Countries
7
Regional context varies; the findings don't.
Employees affected
1M+
The dataset is workforce-scale, not team-scale.
For a PM, this matters because it lets you stop arguing "our sector is different." The shape of what works held across 9 industries.

The 95/5 backdrop — why this study exists

Roughly 95% of generative AI pilots fail to produce measurable financial impact. This is the broader industry context the playbook is responding to. The Stanford team's reframe is that the failure isn't about model quality — it's about treating AI as a technology experiment rather than a production system.

If 95 out of 100 pilots fail, the question to study isn't "what AI feature do we add next?" It's "what did the 5 do that the 95 didn't?" The 51 cases are an attempt to answer that empirically rather than aspirationally.
"Most pilots fail because the model isn't good enough." The playbook's whole argument is that this is rarely true at this point — and where it is true, it's a small slice of the 95%. The dominant failure mode is operational: pilots that never get embedded into a real workflow with real control systems and real ownership.

The central thesis — "the organization, not the model"

One sentence captures the playbook: "The difference was never the AI model. It was always the organization — its readiness, its processes, its leadership, its willingness to change and to fail."

The implication for product strategy is uncomfortable. The most expensive vendor decision PMs make — which foundation model to anchor on — turns out to be one of the least important success drivers. In 42% of the studied cases, the model was interchangeable. What separated winners was workflow redesign, integration depth, operating-model choice, and how leadership showed up on a weekly basis.

Spend less time choosing the model. Spend more time choosing the workflow you'll redesign, the operating model you'll default to, and the executive sponsor you'll trade a weekly cadence with.
If your AI strategy doc spends more pages on "which provider should we use" than on "which workflow are we redesigning," the doc is upside-down relative to what predicts success.

The "weeks vs years" gap

One of the most striking findings: identical use cases took weeks in one company and years in another. The technology was the same. The deployment timeline was set by organizational context — executive sponsorship cadence, infrastructure readiness, end-user willingness, and how aggressively gatekeepers (legal, HR, compliance) blocked progress.

Determines speed
What that looks like in practice
Executive sponsorship cadence
Weekly working sessions vs quarterly status updates.
Infrastructure readiness
Integration paths already built vs greenfield engineering for every connection.
End-user willingness
A pilot population that wants the tool vs one that was assigned it.
Gatekeeper aggression
Legal/HR/compliance running parallel or running blocker.
If you can only diagnose one variable before you forecast a roadmap, diagnose the executive sponsorship cadence. It's the single biggest gating factor between "weeks" and "years."

The 10 findings at a glance

The full playbook organizes its insights into ten findings. Three are unpacked in Guide 02 (the PM playbook); the others are in Guide 03 (operating an AI transformation). Here's the index:

#
Finding
Guide
1
77% of implementation challenges are non-technical (change management, data quality, process redesign)
02
2
61% of successful deployments followed at least one failed attempt
02
3
Organizational context determines speed (weeks vs years for same use case)
01
4
Escalation-based operating models drove 71% median productivity gains
02
5
Active executive involvement (weekly cadence, blocker removal, OKR alignment) drives outcomes
02
6
Resistance comes from legal, HR, compliance — not end users
02
7
Productivity gains don't auto-trigger layoffs — 45% reduced headcount, others redeployed
03
8
Revenue impact from AI remains rare — most projects target cost savings
03
9
Messy data is no longer a blocker — LLMs handle unstructured inputs
03
10
Model choice is increasingly interchangeable — 42% of cases had replaceable models
03
Quiz — Foundations
1. Which of these does NOT belong in Stanford's four-part "successful deployment" bar?
The four criteria are: live in production, used consistently, measurable business value, capable of scaling. Marketing visibility is not on the list.
2. The dataset covered how many organizations and industries?
51 deployments across 41 organizations, 9 industries, 7 countries, 1M+ employees combined. Some orgs had multiple deployments in the dataset.
3. The playbook's central thesis is best summarized as:
"The difference was never the AI model. It was always the organization." Forty-two percent of cases had interchangeable models.
4. Why does the 95% pilot failure rate matter to the playbook's framing?
If most pilots fail, the useful question is what the small minority did differently. The 51 are an empirical answer to that.
5. Identical use cases took weeks in some companies and years in others. The biggest single determinant of that gap was:
Same use case, same technology, vastly different timelines — set by exec cadence, infrastructure readiness, user willingness, and gatekeeper behavior.
6. In what fraction of studied cases was the model interchangeable?
42% of cases had interchangeable models. That's the empirical case for treating model choice as substitutable infrastructure.
7. Your VP says "let's pick the best foundation model first, then design the workflow." Based on the playbook, what's the most accurate response?
The playbook's clearest PM-facing implication. Model lock-in budget should go to workflow, integration, and UX — not to provider choice.
Flashcards — Foundations
01 · Foundations
The Enterprise AI Playbook
tap to reveal →
Stanford Digital Economy Lab report (April 2026, Pereira/Graylin/Brynjolfsson). 51 deployments × 41 orgs × 9 industries × 7 countries. Studies what survived to production.
← tap to flip back
01 · Foundations
The four-part inclusion bar
tap to reveal →
A case counts as "successful" only if it's live in production, used consistently, delivering measurable business value, and capable of scaling. All four — or it's not in the dataset.
← tap to flip back
01 · Foundations
The 95/5 backdrop
tap to reveal →
~95% of generative AI pilots fail to produce measurable financial impact. The Stanford 51 are the reverse-engineered profile of the 5% that worked.
← tap to flip back
01 · Foundations
The central thesis
tap to reveal →
"The difference was never the AI model. It was always the organization." Readiness, processes, leadership, willingness to redesign workflows and absorb failure.
← tap to flip back
01 · Foundations
The "weeks vs years" gap
tap to reveal →
Identical use cases took weeks in some companies, years in others. Set by exec sponsorship cadence, infrastructure readiness, end-user willingness, and gatekeeper behavior.
← tap to flip back
01 · Foundations03 · Advanced
Model interchangeability (42%)
tap to reveal →
In 42% of studied cases, the foundation model was interchangeable without changing outcomes. The empirical case for treating model choice as substitutable infrastructure.
← tap to flip back
01 · Foundations
Stanford Digital Economy Lab
tap to reveal →
Research lab led by Erik Brynjolfsson studying how digital technology is reshaping the economy. Publisher of the Enterprise AI Playbook report.
← tap to flip back
01 · Foundations
Erik Brynjolfsson
tap to reveal →
Stanford economist, co-author of the playbook. Leading researcher on technology and economic productivity (Second Machine Age, Race Against the Machine).
← tap to flip back
01 · Foundations
Experiment vs production framing
tap to reveal →
The most-cited failure pattern: teams treat AI projects as side-pocket experiments rather than production systems with control, integration, and operational ownership.
← tap to flip back
01 · Foundations02 · Applied
Workflow-redesign separator (55% vs 20%)
tap to reveal →
McKinsey signal cited in the playbook: 55% of high performers redesigned workflows around AI vs only 20% of other companies. The largest single separator.
← tap to flip back
Library · Next: 02 · Applied →
Full Quiz · All Flashcards