Skip to content
transformative-ai
MethodologyDiagnosticEngagementInsightsAbout
Explore an Engagement →
transformative-ai
MethodologyDiagnosticEngagementInsightsAbout
Explore an Engagement →
Insights/Essays
ESSAY

From pilot to production: a 90-day AI roadmap.

A pilot proves a model can work; production proves the firm can absorb it, and that second proof needs a different plan.

Published
October 2026
Reading time
9 minutes
Author
Transformative AI
inX

An AI pilot to production roadmap is a time-boxed plan for moving an AI system out of a controlled trial and into a live business process that depends on its output. The pilot proves a model can do the work. The roadmap proves the firm can absorb it.

Most plans only budget for the first proof. What takes the time is everything the pilot was allowed to skip: whose data it reads, who may act on what it says, which process must consume it, and how anyone will know, six months in, whether it is still right.

Why most enterprise AI projects fail names the three ways pilots die: the orphaned pilot, the contested boundary, and the human in the critical path. This essay asks the practical follow-up. Given one pilot that works, what do the next ninety days look like?

§ 01 / THE STALLWhy pilots succeed and then stall

A pilot answers one question: can a model do this task well enough? It is scoped to make the answer easy: clean extract, friendly sponsor, a metric chosen in advance, a reviewer checking every output. Under those conditions the answer is usually yes.

Production asks whether the firm can run the decision differently, and nothing about a successful pilot answers that. The public evidence on how often it goes unanswered is consistent, if imprecise.

MIT NANDA's The GenAI Divide: State of AI in Business 2025 reports that 95 percent of organizations are getting zero return from generative AI, while just 5 percent of integrated pilots are "extracting millions in value" and the vast majority remain stuck with no measurable P&L impact. Read the method before quoting the number. It rests on a review of over 300 public initiatives, interviews with representatives of 52 organizations, and 153 survey responses from senior leaders, and the authors themselves describe the pilot-to-production figures as "directionally accurate based on individual interviews rather than official company reporting." Treat it as a strong signal, not a census.

The more useful line in that report is its diagnosis. Enterprise tools, it finds, mostly fail "due to brittle workflows, lack of contextual learning, and misalignment with day-to-day operations." None of those is a model problem. Gartner reached a similar place from a different direction, predicting in July 2024 that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025 (a prediction whose window has now closed), citing poor data quality, inadequate risk controls, escalating costs or unclear business value.

On the other side of the ledger, McKinsey's March 2025 State of AI report (survey fielded July 2024) found that out of 25 attributes tested, "the redesign of workflows has the biggest effect on an organization's ability to see EBIT impact from its use of gen AI." Only 21 percent of respondents whose organizations use generative AI said they had fundamentally redesigned at least some workflows.

A pilot asks whether the model can do the work. Production asks whether the firm can run the decision differently.

§ 02 / READINESSThe integration readiness checks

Five questions decide whether a pilot has a path into the operating loop. Each maps to a way the firm can reject the system, and each has a binary answer.

  1. Data contracts. Is the data the system consumes published under a versioned contract with a named owner, a schema, and a freshness commitment? Or does the pilot read a one-time extract someone pulled by hand? Gartner predicts that through 2026 organizations will abandon 60 percent of AI projects unsupported by AI-ready data. An uncontracted feed is one of the most common forms of not-ready.
  2. Decision rights. Who is authorized to act on the output, and is that written down? If the answer is "the committee," the system will produce recommendations that wait for a meeting. A decision with no owner is how the contested boundary begins.
  3. Who consumes the output. Name the downstream process that is obligated to use the result. Not interested in it, obligated. An output with no committed consumer is the orphaned pilot, however good the model.
  4. The human on the critical path. Where does a person review, and is that review on every case or only on exceptions? Review on every case turns a fast decision into a queue, and a queue means the loop never closes fast enough to learn.
  5. Evaluation. Is there a standing test suite that replays real traffic against the system and scores it on the business outcome, not only on model accuracy? Without one, every model change is a fresh risk review, and the system ossifies at its first version.
THE TESTIf you cannot answer all five with a name, a document, or a running job, the system is still a pilot, whatever the deployment dashboard says.

§ 03 / WEEKS 1 TO 4Weeks 1 to 4: assess

The first month produces no production code. It produces a choice of what to build and a written definition of done. That sounds slow. It is the fastest route, because every unresolved readiness check becomes a stall later, when it is far more expensive to fix.

Week 1: choose one decision. Not a use case, not a capability, a decision: a specific, recurring judgment the business makes, with an owner and an outcome that can be observed. Rank candidates on value, frequency, and how soon the outcome is visible. A decision whose result shows up in days beats a bigger one whose result shows up in a year, because the short loop can learn.

Week 2: map the current loop. Trace the decision as it runs today: where inputs come from, who touches them, what action follows, and whether the outcome is ever connected back to the decision. The honest drawing is usually an open loop: decisions execute, outcomes land in another system, and nothing joins them.

Week 3: run the five checks. Record the current state and the gap for each. Most gaps are organizational: a missing data contract is usually a missing agreement between two teams. Write each gap down with the person who can close it.

Week 4: write the exit criteria and the plan. Define what "in production" will mean for this decision (see § 06), set the business metric and its baseline, and sequence the build. End with a go or no-go. A no-go is legitimate: choosing another decision is cheaper than forcing an unready one.

WORKED EXAMPLEConsider an illustrative mid-market industrial distributor whose credit team reviews every order that trips a credit hold. The decision is release or hold. The outcome, whether the customer paid on time, is visible within weeks but lives in the receivables system, never joined to the hold decision. The pilot scored well on a historical extract. Week 3 finds three gaps: the order feed has no contract, the authority to auto-release is unwritten, and every recommendation still routes to a human. The plan comes from those gaps, not from the model.

§ 04 / WEEKS 5 TO 13Weeks 5 to 13: build one closed loop

The next nine weeks build one closed loop: a decision wired to an action, with the outcome fed back as a signal the system and its owners can learn from. In the methodology, closed loops are the first of the Five Pillars, the physics the others depend on.

Weeks 5 and 6: contract the inputs. Publish the data the decision needs once, under a versioned contract with an owner and a freshness SLA. This is the smallest version of the publish-subscribe data layer the AI Factory depends on: build the slice this loop needs, not the whole layer.

Weeks 7 and 8: wire decision to action. Connect the output to the process obligated to consume it, under the decision rights agreed in week 3. Move human review to the exceptions: low confidence, high value, or outside policy. The aim is not to remove people but to put them where their judgment changes outcomes.

Weeks 9 and 10: stand up evaluation and run in shadow. Build the standing evaluation suite, then run the system in shadow beside the incumbent process. Both make the decision; only the incumbent acts. Compare them on the business metric, scored against live traffic rather than the pilot's clean extract.

Weeks 11 and 12: cut over with a rollback. Move a defined share of volume to the new loop, with a tested path back to the old process, and widen it as the metric holds.

Week 13: close the loop and hand it over. Join each outcome to the decision that produced it, so results are visible per decision, not as quarterly averages. Hand operation to the named owner.

One working loop is worth more than five promising pilots. Once it exists, the architecture is precedent rather than proposal.

§ 05 / THE TEMPLATEThe roadmap on one page

Copy this as a template. Replace every generic noun with your decision, dataset, team, and metric. Each phase ends at a gate, and the gate is a written artifact, not a demo.

Phase 1. Assess (weeks 1 to 4)

  1. Week 1: shortlist recurring decisions; choose one with an owner, a visible outcome, and a short feedback cycle.
  2. Week 2: map the current loop from input to outcome; mark where the loop is open.
  3. Week 3: run the five readiness checks (data contract, decision rights, committed consumer, human on the critical path, evaluation); assign each gap an owner.
  4. Week 4: write exit criteria, the business metric and its baseline, and the build sequence.
  5. Gate: written go or no-go, signed by the decision owner.

Phase 2. Build one closed loop (weeks 5 to 13)

  1. Weeks 5 and 6: publish the input data under a versioned contract with an owner and freshness SLA.
  2. Weeks 7 and 8: wire the output to the obligated consumer; move human review to exceptions.
  3. Weeks 9 and 10: build the evaluation suite on replayed traffic; run in shadow against the incumbent.
  4. Weeks 11 and 12: staged cutover with a tested rollback path.
  5. Week 13: join outcomes to decisions; hand operation to the named owner.
  6. Gate: the production criteria in § 06 are met and documented.

Phase 3. Repeat

  1. Choose the second decision, preferably one that reuses the contracted data from the first.
  2. Run it through the same gates. It should be faster, because the hard parts are now precedent.

§ 06 / THE DEFINITIONWhat "in production" actually means

"In production" usually gets used to mean deployed. That is a hosting fact, not a business one. A system is in production when the business would notice, within days, if it stopped. Concretely, all of the following are true. Its inputs arrive under a contract that someone owns and that breaks loudly when violated. Its outputs drive an action without a meeting. People review exceptions, not every case. A standing evaluation suite can score a candidate model change against real traffic, so a better model is a configuration change rather than a project. Each outcome is joined back to the decision that produced it. And a business metric that belongs to a P&L, with a baseline set before the build, has moved.

That last test is the one most systems never face. McKinsey's survey found that fewer than one in five respondents said their organizations track KPIs for gen AI solutions, and identified KPI tracking as the adoption practice, of the 12 tested, with the most impact on the bottom line. A loop with no metric is not yet a loop.

DEFINITIONIn production: a decision the business runs on, fed by contracted data, acted on without a meeting, reviewed by exception, continuously evaluated, and measured against a P&L metric it has demonstrably moved.

This roadmap is the shape of the work, not the fitted version. The four-week Assessment is the fitted version of weeks 1 to 4: a fixed-fee assessment of the firm's current architecture whose deliverable is a written report identifying the three highest-leverage interventions and a 90-day execution roadmap for the build that follows. Where more than one loop must close, the architectural engagement rebuilds one operating system to AI-Native standard over 90 to 180 days, and the embedded partnership operates the transformation end to end over twelve months or more.

For a smaller first step, the free Maturity Diagnostic takes eight questions and scores your firm on six axes (what each axis measures). If the loop you have in mind will be run by agents, the agentic AI operating model covers the structures it needs. Either way, start with one decision. The firms that get AI into production are not the ones with the most pilots. They are the ones that closed the first loop.

END

§ FAQ / READER QUESTIONSQuestions readers ask

What is an AI pilot to production roadmap?

It is a time-boxed plan for moving an AI system from a controlled pilot into a live business process that depends on it. A useful one spends its first weeks on integration readiness (data contracts, decision rights, the consumer of the output, human review, evaluation) and the rest on building one closed loop, where a decision, its action, and its measured outcome are connected and owned.

What should a 90-day AI roadmap include?

Two phases. Weeks 1 to 4 assess: pick one decision, map its current loop, run the five readiness checks, and write the exit criteria. Weeks 5 to 13 build: publish the data contract, wire the decision to an action, stand up evaluation, shadow the incumbent process, then cut over with a rollback path. Each phase ends with a written gate, not a demo.

What belongs on an AI POC to production checklist?

A versioned data contract with an owner and a freshness SLA; a named person with the authority to act on the output; a downstream process obligated to consume it; human review placed on exceptions rather than every case; an evaluation suite that replays real traffic; a business metric owned by a P&L; and a rollback path that has actually been tested. If any item is missing, the system is still a pilot.

How do you use an AI transformation roadmap template?

Treat it as a sequence of gates, not a calendar. Copy the phases, replace each generic item with the specific decision, dataset, team and metric in your firm, and refuse to advance until the gate is met in writing. The template is only useful for one loop at a time; running five loops through it in parallel recreates the conditions that stall pilots.

Why do most AI pilots fail to reach production?

Public research points to integration rather than model quality. MIT NANDA's 2025 report finds that enterprise AI tools mostly fail due to brittle workflows, lack of contextual learning and misalignment with daily operations, and Gartner cites poor data quality, inadequate risk controls, escalating costs and unclear business value. In practice those causes show up as unowned data, unclear decision rights and a human review step that keeps the loop open.

Transformative AI

AN ENGINEERING-LED PRACTICE

We rebuild enterprise operating runtime for the AI-Native era: closing loops, enforcing data contracts, and moving human review off the critical path. The deliverable is not a deck. It's a redesigned firm.

CONTINUE READING

ESSAY · RUNTIME PAPERS No. 4MAY 2026

Why most enterprise AI projects fail before they ship.

The pattern is consistent: pilots succeed, production deploys, the system never integrates. The diagnosis is architectural, not technical.

11 MIN→
ESSAYOCT 2026

What is an agentic AI operating model?

An agentic AI operating model is the architecture, evaluation discipline, ownership, and governance that let agents do accountable work inside a firm. Without it, agents stay pilots.

11 MIN→
THE DELIVERABLE IS NOT A DECK

Find out where your runtime rejects the work.

Start the Conversation →
transformative-ai

An engineering-led practice for enterprises rebuilding the machine the company runs on, for the era when algorithms and networks carry the load.

AN ENGINEERING-LED PRACTICE

The Site

  • Methodology
  • Diagnostic
  • Engagement
  • Insights
  • About

Contact

  • Explore an Engagement
  • Contact

Elsewhere

  • LinkedInSoon
  • X / TwitterSoon
  • SubstackSoon
  • RSS
© 2026 TRANSFORMATIVE AIBUILT FOR THE THIRD STATE