Field Note · Pilot-to-Production Gap

Why AI Hasn't Improved Productivity: What a 6,000-Executive Survey Actually Reveals

Nine in ten executives report no measurable AI impact in three years. The same executives forecast gains ahead. That gap is not hypocrisy, it is an operational tell.

8 min read Published August 13, 2026

DefinitionThe pilot-to-production gap is the failure of most AI investments to produce measurable business results. It is an operational failure, not a technical one. Companies adopt the tools, run pilots, and see encouraging demos, but the systems never integrate into how work actually moves through the business.

In February 2026, the National Bureau of Economic Research published Working Paper 34836, a survey of nearly 6,000 senior executives across the United States, United Kingdom, Germany, and Australia. The headline finding: 89% reported that AI had no measurable effect on their firm's productivity or employment over the prior three years. The same executives forecast meaningful gains ahead, a 1.4% productivity improvement on average.

That gap between experience and expectation is not hypocrisy. It is an operational tell.

The pilot-to-production gap is the failure of most AI investments to produce measurable business results. Unlike a technical failure, it is an operational one: companies adopt the tools, run pilots, and see encouraging demos, but the systems never integrate into how work actually moves through the business. According to research compiled by Vovance in March 2026, only 26% of AI pilots reach sustained production use. The executives in the NBER survey are not outliers. They are the 74%.

89%
Report no measurable AI impact on their firm
26%
Of AI pilots reach sustained production
+1.4%
Productivity gain the same executives forecast

Sources: NBER Working Paper 34836, February 2026. Vovance synthesis, March 2026.

Why the Data Is Not Surprising

If you have watched an organization try to adopt AI over the past three years, the NBER finding matches what you have seen.

The pattern is consistent: a company pays for AI tools, employees use them individually for drafting emails and summarizing documents, productivity improves at the margins for those individuals, and when a C-suite executive asks what changed in how the business operates, the answer is: not much.

That outcome is the predictable result of treating AI as a tool rather than a system. Tools are adopted and systems are built. The difference is whether the AI is integrated into how work moves through the organization, or whether it sits alongside the existing workflow, available for individuals to use when they remember.

The NBER data captures this distinction in its own framing: executives see no measurable firm level impact. That is the systems question. Individual productivity gains from using AI writing tools are real, but they do not aggregate to the business results executives are waiting to see unless the workflows are designed to capture them.

Tools are adopted. Systems are built. The 89% adopted.

Three Operational Failure Modes the Data Points To

Based on the NBER findings and the pilot-to-production gap research, the executives not seeing results tend to share one or more of the following operational conditions. These are the same patterns that surface in a Radiant Work operations audit.

1 · Operational Foundation

Documented workflows, one source of truth, consistent data.

The substance an AI system has to run on. Without it, there is nothing for an agent to be accurate about.

Skip it, and AI runs on disorganized operations, faster.
Workflow integration depends on it
2 · Workflow Integration

AI embedded in the process, not available alongside it.

The work is redesigned so the AI sits inside the path a job actually takes through the business.

Deploy without it, and individual gains never aggregate to the firm.
Measurement depends on it
3 · Specific Success Criteria

One measurable outcome, a 60 to 90 day window.

Hours reclaimed, error rate, cycle time. Named before the build, checked after it.

Leave it diffuse, and the result is indistinguishable from no result.
Business impact depends on it
4 · Business-Level Impact

The number an executive can actually see.

What the 11% report and the 89% do not. It is downstream of all three prior steps, in order.

Expect it without them, and you get the NBER outcome.

Each step is a prerequisite for the next. Skipping any node produces the 89% result.

They Deployed AI in Silos

A firm that gives its team a ChatGPT subscription has adopted AI. If the project management workflow, the vendor coordination process, the client communication cadence, and the proposal drafting process all remain unchanged, the AI has not entered the operating system of the business. It has entered the browser tabs of several employees.

Siloed AI use produces individual gains. Business level gains require that AI be integrated and automated at the workflow level, which means redesigning how work moves, not just adding a tool to an existing process that was designed before the tool existed.

This is also where the expectations math breaks down. An executive who expects AI to improve firm level productivity but has only deployed it as an individual tool is measuring the wrong thing. The gap between forecast and experience in the NBER data is partly a gap between what was deployed and what would be required to produce the result they are expecting.

They Expected Transformation Before Proving Small Wins

The NBER finding that executives expect future gains while reporting no past gains suggests a pattern familiar in change management: organizations launched at scale before proving value at micro scale. The expectation was transformation; the implementation skipped the validation steps that would have shown whether the specific application worked.

The businesses that get measurable results from AI tend to start smaller and prove faster. A single repeatable step in a workflow, automated and measured over 60 days. A single agent, with a narrow job and a clear success criterion. The compound effect of small validated wins is more reliable than a broad transformation program that never reaches production.

The instinct to go big first is understandable given the scale of AI's promise. The problem is that a big deployment with a diffuse success definition produces a diffuse result, which is indistinguishable from no result at all.

Their Infrastructure Was Not Ready for the Expectation

The NBER survey captures a striking prediction: executives expect 1.4% productivity gains ahead. For a ten person business, that is roughly a half hour of additional productive time per person per day. A modest, achievable expectation.

But achieving even a modest expectation requires that the business has the operational infrastructure to support it:

  1. a consistent source of truth for business information,
  2. workflows that are documented and repeatable,
  3. and systems that integrate rather than silo data.

Many businesses that have not seen results have invested in AI tools without investing in the operational foundation those tools require.

Bolting AI onto disorganized operations does not produce measurable gains. It produces a faster way to generate disorganized outputs. The executives forecasting 1.4% gains next year are implicitly forecasting that their operations will have matured enough to support it. That is the real variable.

What the 11% Are Doing Differently

The NBER study identified a minority of executives who do report measurable impact. The research does not isolate the cause, but the pattern that shows up consistently in the AI implementation literature is the same: the organizations seeing results addressed the operational prerequisites before expanding AI scope.

That means a single source of truth for business information. Documented workflows that AI can integrate with rather than work around. Clear ownership of who manages AI outputs and corrects the system when it drifts. More importantly, moving from discussing work within an AI chat window to directing it to a tangible business output.

It also means starting with the workflows that have the clearest success criteria. Not "let's use AI everywhere" but "let's use AI on this specific step and measure what changes." That specificity is what separates a deployed system from a pilot that never reached production.

The gap between the 89% and the 11% is an operational gap, not a technology gap. The models available to both groups are the same. What differs is the infrastructure the models are operating within.

Related Questions

Why do most AI pilots fail to reach production?

According to research compiled by Vovance in March 2026, only 26% of AI pilots reach sustained production use. The primary failure modes are missing context, missing integration, and missing governance, operational problems rather than technology problems. Most pilots fail because the operational foundation was not in place before the AI was deployed.

How long does it take to see measurable AI impact?

Measurable impact at the individual level can appear within weeks when a tool addresses a specific repetitive task. Business level impact, the kind that shows up in productivity or revenue metrics, typically requires 60 to 90 days of a redesigned workflow operating with AI integrated into it, not just available to individuals.

What should executives do if AI hasn't improved their business?

The first step is an honest assessment of where AI is deployed: is it integrated into a workflow, or is it available as a tool? If AI is not integrated, the question is not "what different AI should we try" but "what operational change would allow AI to produce results?" That assessment is the starting point.

Is the AI productivity gap a problem with the technology?

No. The NBER finding, and the broader pilot-to-production gap research, consistently point to operational causes. The models are capable of the tasks most businesses need. The gap is in how those models are integrated, what context they operate with, and whether the workflows around them are designed to produce business results.

What is the relationship between operations and AI results?

Operations are what AI runs on. An AI agent is only as useful as the information it can access, the workflow it is integrated into, and the governance that keeps it accurate over time. Businesses that improve their operations before deploying AI at scale consistently report better results than those that deploy AI first and expect operations to follow.

The Work Behind the Work

The gap between the 89% and the 11% is operational. It is also fixable.

Take the first step toward a business that runs with clarity and momentum.

For Deeper Context

  1. National Bureau of Economic Research, Working Paper 34836 (February 2026). Survey of approximately 6,000 senior executives across the US, UK, Germany, and Australia. nber.org/papers/w34836
  2. Vovance synthesis of a March 2026 enterprise AI survey, n=650. Why Most Enterprise AI Pilots Never Make It to Production, and What the Survivors Did Differently. medium.com