Skip to content
transformative-ai
MethodologyDiagnosticEngagementInsightsAbout
Explore an Engagement →
transformative-ai
MethodologyDiagnosticEngagementInsightsAbout
Explore an Engagement →
Insights/Essays
ESSAY

What an AI maturity assessment should actually measure.

Most maturity models inventory capabilities; the useful question is what the operating machine is made of, and which of six axes binds it first.

Published
October 2026
Reading time
10 minutes
Author
Transformative AI
inX

An AI maturity assessment is a structured measurement of how far artificial intelligence has moved from side project to operating machinery inside a firm. A useful one scores what the business does today (where signal goes, where judgment lives, what initiates work, who can change a model) rather than what it plans, buys or intends.

That distinction sounds pedantic. It is the whole difference between an assessment that changes a roadmap and one that decorates it. Asked to rate their own AI maturity, firms tend to describe their strategy, their platform purchases and their pilot count. None of those tell you whether the machine has changed. The questions that do are smaller, more specific and harder to answer aspirationally.

§ 01 / DEFINITIONSWhat maturity measures, and readiness doesn't

Readiness and maturity are routinely used as synonyms. They measure different things.

Readiness is an inventory of preconditions. Cisco's AI Readiness Index, for example, assessed organizations across six pillars: strategy, infrastructure, data, talent, governance and culture, and found that only 13% of the large companies surveyed were fully ready to capture AI's potential. That is a useful number. It answers the question "could this firm do it?"

Maturity answers a different question: "is this firm doing it?" A company can own a modern data platform, a signed AI policy and a capable data-science team, and still route every below-floor pricing exception through an email chain. It is ready. It is not mature. The inverse is hard to construct, which is why readiness is a necessary input and a poor proxy for the output.

DEFINITIONReadiness measures whether the preconditions for AI exist. Maturity measures whether AI is load-bearing in how the firm actually operates: whether signal, decisions and work move through models or around them.

The practical consequence is in how the questions are phrased. A readiness question asks whether you have a data strategy. A maturity question asks what happens to a customer interaction signal after it is captured. The first invites a slide. The second has a factual answer, and the answer is usually less flattering.

§ 02 / THE LANDSCAPEHow the established models compare

Three frameworks are worth holding side by side, because each is built to answer a different question.

Gartner's AI maturity model. Gartner groups organizations into five stages: Foundational, Emerging, Operational, Scaled and Transformational. Its assessment scores seven core pillars (strategy, data, governance, engineering, operating model, culture and AI product or value) and compares current maturity against target maturity for each, so the largest gaps can be prioritized into a roadmap. Its strength is breadth. It is, at heart, a capability inventory: a disciplined way to see what a firm has and lacks across the functions that touch AI.

The MIT CISR Enterprise AI Maturity Model. MIT CISR defines four stages: Experiment and Prepare, Build Pilots and Capabilities, Develop AI Ways of Working, and Become AI Future Ready. Its distinctive move is tying stage to money. Based on a 2022 survey of 721 companies, most enterprises were in the first two stages and had financial performance below industry average, while enterprises in stages three and four had financial performance well above industry average. A 2025 follow-up found that the greatest financial impact is achieved in progressing from stage 2 to stage 3, the step from piloting AI to scaling it. Respondents were staged by a measure of AI effectiveness across operations, customer experience and ecosystems. Where Gartner inventories capability, MIT CISR measures effectiveness.

The AI-Native spectrum. The spectrum in our methodology has three states. AI-Enabled firms apply AI to specific functions to make existing processes smarter: a turbocharger on an internal combustion engine. AI-First firms redesign the core around AI, but on legacy foundations: a new dashboard and control system, old engine. AI-Native firms are conceived and built on AI, to the point that without AI the business model ceases to exist: a new electric engine, reimagined from the chassis.

The spectrum asks a third question that neither of the others asks directly: what is the operating machine made of? Not how many capabilities the firm holds, and not how effective its AI is in aggregate, but whether AI is structural or attached. A firm can score respectably on a capability inventory and still be running an AI-Enabled machine, because every capability it owns has been bolted onto workflows designed for human throughput.

For a mid-market firm, the comparison has a practical edge. Gartner's toolkit page presents the full in-depth report as a client deliverable, and MIT CISR's model is published as research for locating yourself conceptually. Both are valuable. Neither is designed to tell a leaner leadership team, in an afternoon, which single constraint to remove first. That is the job a short, behavioral AI maturity assessment can do, provided its questions are about the machine and not the plan.

§ 03 / THE INSTRUMENTThe six axes defined

The free Maturity Diagnostic scores a firm on six axes using eight questions. Each question offers four answers, scored from zero (legacy) to three (AI-native), and the answers are shuffled every session so position never signals the "right" one. Two axes, Workflow and Strategy, are compound: each takes two questions, averaged. Every axis is normalized to a percentage. The overall score is the unweighted mean of all six.

1. Data. Operational signal, captured and put to work. The question: a customer interacts with your product; where does that signal go? The answers run from a report someone reads next quarter, through a weekly dashboard and an on-demand warehouse, to models that retrain and act on it automatically. This is the raw-material function of the AI Factory. If signal stops at a human reader, every downstream model is starved at the source.

2. Decisions. Where judgment lives: inbox or model. The question: a deal is proposed below the approved price floor; what happens next? At the bottom, it routes through email until someone signs off. At the top, a model prices it and logs the reasoning in seconds. Between them sit the deal desk and the rule that auto-approves standard cases. The decision layer is where latency and inconsistency compound, and it is the slowest layer to retrofit.

3. Workflow. What initiates the work, and what catches drift. Two questions. First, how much operating work begins with a human noticing something: almost all of it, or little, with events triggering work and humans handling exceptions? Second, when a process quietly degrades, how is it caught: in a quarterly review weeks later, at a manual threshold, by monitoring that pages a team, or by models that detect, attribute and intervene? Together they measure whether the firm runs closed loops or open ones.

4. Architecture. How load-bearing AI is to the operating machine. The question is blunt: could the business run for a day with its AI systems switched off? "Easily" means AI is a side experiment. "No: without AI the model ceases to exist" means AI is the runtime. This is the axis that most directly separates the three states of the spectrum.

5. Talent. Who can ship a change to a model. The question: who can ship a change to a model running in production? The answers move from no one (the firm runs no models of its own), to a small separate data-science team, to engineers with data-science review, to product teams for whom models are part of the stack. A separate priesthood caps how fast capability reaches the product, however good the models are.

6. Strategy. Where the roadmap's center of gravity sits. Two questions. How does leadership treat AI: as a cost line to control, a set of point tools, a capability to build toward, or the architecture the firm is rebuilt around? And is the roadmap defending the current model, bolting AI onto what exists, redesigning core workflows, or compounding advantage that widens over time?

The overall score maps to a band. Below 34% reads as AI-Enabled. Below 67% reads as AI-First. Above that reads as AI-Native. The readout shows the band, not the underlying number.

Five of the six axes ask what happens, not what is planned. That is deliberate. Behavior is harder to round up.

The weighting toward operations over technology is not arbitrary. BCG's 2024 survey of 1,000 executives found that around 70% of AI implementation challenges stem from people- and process-related issues, 20% from technology and only 10% from algorithms. An instrument that spent most of its questions on model quality would be measuring the smallest part of the problem.

§ 04 / INTERPRETATIONReading your radar

The band is the least informative thing the diagnostic returns. Because the overall score is a simple mean, two very different firms can land in the same place.

WORKED EXAMPLEConsider an illustrative firm whose customer signal feeds models that retrain automatically (Data 100%), whose below-floor deals go to a deal desk (Decisions 33%), whose work mostly starts with a human but whose monitoring pages a team (Workflow 50%), whose workflows would stall without AI (Architecture 67%), whose models are owned by a small separate team (Talent 33%), and whose leadership is building toward AI with a compounding roadmap (Strategy 83%). The mean is 61%: AI-First. A second firm that picks the third answer on every question scores a flat 66.7%, a fraction under the 67% AI-Native threshold, and is also AI-First. Same band. Entirely different problems.

The shape is the finding. A few patterns are worth reading for.

Strategy high, Architecture and Decisions low. Ambition is running ahead of the machine. Leadership describes AI as the architecture; the business would run fine with it switched off, and judgment still lives in inboxes. This is how a strategy deck becomes an open loop: the intent is real, the operating path from intent to outcome does not exist yet.

Data high, Decisions low. The firm captures signal well and then hands it to a human to act on. The loop is closed at the front and open at the back. Models produce predictions no process is obligated to consume, which is exactly the orphaned-pilot pattern described in why most enterprise AI projects fail.

Talent low, anything else high. Every other axis is rate-limited by this one. If only a small, separate team can touch a production model, improvements queue behind that team no matter how good the data or how clear the strategy.

Workflow low. This one has outside evidence behind it: in McKinsey's survey research, workflow redesign is the attribute most associated with seeing bottom-line impact from generative AI, and few organizations have done it (we cover the numbers in the pilot-to-production roadmap). A low Workflow score is a direct read on the attribute most associated with value.

Which axis binds first? The diagnostic answers mechanically: it names your strongest axis, the axis where the gap hurts most, and your three lowest-scoring axes as the three highest-leverage moves. That ranking is the right starting point. Read it with one structural rule in mind: a closed loop runs at the speed of its slowest stage. Signal capture (Data), judgment (Decisions) and initiation (Workflow) are stages of the same loop, so the lowest of those three is usually the constraint on the other two. Talent and Strategy set the rate at which any constraint can be removed. Architecture is the result: it rises when the loops close, not before.

THE TESTIf you raised your strongest axis to 100% tomorrow, would anything in the business move faster? If the answer is no, your strongest axis is not where the work is.

§ 05 / ACTIONWhat to do with the score

A maturity score is not a grade. It is a pointer to the next loop worth closing.

Start with the axis the diagnostic names as the gap that hurts most, and find one concrete process where that gap bites. If it is Decisions, that might be the below-floor pricing path itself. Then make the loop real, the way the methodology describes it: the data the decision consumes published under a contract, the decision wired to an action that happens without a meeting, and the outcome fed back as a signal the system learns from. One closed loop, with an owner and a business metric, does more for a maturity score than a portfolio of pilots, because it changes the answer to the questions rather than the description of the answers.

Then run the assessment again. The movement you are looking for is not a higher band. It is the lowest axis rising and the shape getting rounder. MIT CISR's research locates the largest financial gain at the move from piloting AI to scaling it as a way of working. In spectrum terms, that is the work of turning isolated AI-Enabled wins into AI-First workflows, and it is earned one loop at a time.

IN PRACTICEA self-scored diagnostic is a hypothesis about your firm. The way to test it is to read the firm's data contracts and decision rights directly: who owns each input, who can approve each exception, and how long each loop actually takes to turn.

That test is what the four-week Assessment engagement does. It reads the current architecture directly and ends in a written report naming the three highest-leverage interventions and a 90-day execution roadmap. The details are on the engagement page.

Start smaller. The Maturity Diagnostic is free, takes about two minutes, asks eight questions and does not ask for your email. It returns your band, your strongest axis, the axis where the gap hurts most and your three highest-leverage moves. Take the eight-question Maturity Diagnostic, look hard at the shape, and then decide which loop to close first.

END

§ FAQ / READER QUESTIONSQuestions readers ask

What is the difference between AI readiness and AI maturity?

AI readiness measures whether the preconditions for AI exist: strategy, infrastructure, data, talent, governance and culture. AI maturity measures how AI actually runs the business today: where operational signal goes, who makes decisions, what initiates work and whether the firm depends on its models. A firm can own modern infrastructure and still route every exception through email. Readiness is potential; maturity is observed behavior.

What questions should an AI maturity assessment ask?

Good AI maturity assessment questions ask what happens, not what is planned. Where does a customer interaction signal go? What happens when a deal is proposed below the price floor? How much work begins with a human noticing something? Could the business run for a day with AI switched off? Who can ship a change to a production model? Behavioral questions are harder to answer aspirationally, so the score reflects the operating machine rather than the strategy deck.

How do the Gartner and MIT CISR AI maturity models compare?

Gartner's AI maturity model groups organizations into five stages, from Foundational to Transformational, and scores seven pillars including strategy, data, governance, engineering, operating model and culture. MIT CISR's Enterprise AI Maturity Model uses four stages and links them to financial performance, finding that firms in stages three and four performed well above industry average. Gartner inventories capability; MIT CISR measures effectiveness. Neither asks directly whether AI is load-bearing in the operating architecture.

Is an AI maturity assessment useful for mid-market firms?

Yes, and arguably more so, because mid-market firms cannot fund parallel transformation programs and need to know which single constraint to remove first. The useful assessment for a mid-market firm is short, behavioral and pointed at the binding axis rather than a long capability inventory. A two-minute diagnostic that names the weakest axis, followed by a focused review of data contracts and decision rights, fits the budget and the attention span of a leaner leadership team.

How is the Transformative AI Maturity Diagnostic scored?

The Maturity Diagnostic asks eight questions across six axes: Data, Decisions, Workflow, Architecture, Talent and Strategy. Workflow and Strategy each take two questions. Every answer scores from zero to three, each axis is averaged and normalized to a percentage, and the overall score is the unweighted mean of the six axes. Below 34 percent reads as AI-Enabled, below 67 percent as AI-First, and above that as AI-Native.

Transformative AI

AN ENGINEERING-LED PRACTICE

We rebuild enterprise operating runtime for the AI-Native era: closing loops, enforcing data contracts, and moving human review off the critical path. The deliverable is not a deck. It's a redesigned firm.

CONTINUE READING

ESSAYOCT 2026

From pilot to production: a 90-day AI roadmap.

Most AI pilots stall at the boundary of the firm, not the limit of the model. Here is a 90-day roadmap that spends four weeks on readiness and nine on closing one loop.

9 MIN→
ESSAY · RUNTIME PAPERS No. 4MAY 2026

Why most enterprise AI projects fail before they ship.

The pattern is consistent: pilots succeed, production deploys, the system never integrates. The diagnosis is architectural, not technical.

11 MIN→
THE DELIVERABLE IS NOT A DECK

Find out where your runtime rejects the work.

Start the Conversation →
transformative-ai

An engineering-led practice for enterprises rebuilding the machine the company runs on, for the era when algorithms and networks carry the load.

AN ENGINEERING-LED PRACTICE

The Site

  • Methodology
  • Diagnostic
  • Engagement
  • Insights
  • About

Contact

  • Explore an Engagement
  • Contact

Elsewhere

  • LinkedInSoon
  • X / TwitterSoon
  • SubstackSoon
  • RSS
© 2026 TRANSFORMATIVE AIBUILT FOR THE THIRD STATE