Skip to content

From Pilot Factory to Product Factory: How CIOs Build a Repeatable AI Delivery Engine

Published: at 12:50 AMSuggest Changes

There is a slide that appears in most enterprise AI reviews. It shows the portfolio: twenty-two use cases in discovery, nine in pilot, four “scaling”, one with a signed business case. The count has grown since last quarter, which is presented as progress. Then someone asks which of the twenty-two the organisation would defend if the budget were halved, and the room goes quiet.

The silence is the diagnosis. Most organisations have built a pilot factory: a reliable machine for producing experiments, demonstrations and steering-committee material. Very few have built a product factory: a machine for turning a small number of those experiments into services that run in production, carry a named owner, cost a known amount per task and get retired when they stop earning their keep.

The distinction matters because the two machines are optimised for different things. A pilot factory is rewarded for learning, optionality and visible momentum. A product factory is rewarded for operated reliability, unit economics and evidence. Improving the first will not produce the second. CIOs who assume a healthy pilot portfolio will “graduate” into production are waiting for a conversion that nobody owns.

The conversion gap is the whole problem

The evidence on this is now consistent enough to stop arguing about.

MIT’s Project NANDA analysed 300 public deployments, 150 executive interviews and a survey of 350 employees, and concluded that 95% of organisations were getting no measurable return from generative AI, with only about 5% of custom enterprise tools reaching production. The more useful finding was about time: mid-market firms moved from pilot to full implementation in roughly 90 days, while large enterprises took nine months or longer. Enterprises ran the most pilots and converted the fewest.

Gartner’s September 2026 research, based on a survey of 1,303 functional leaders, found that only 22% of organisations had successfully scaled AI across multiple business units. In the same research, 11% of functions could not say what they had spent on AI in 2025 at all.

Read those together and the problem is not model quality, data science talent or executive appetite. It is that the organisation has no production line. Pilots are cheap to start and have no defined end. Production has an owner, a cost line, a control set and a retirement condition. The distance between the two is a design gap, and it sits with the CIO.

Publish the intake criteria and the takt time

A factory is defined by what it refuses. Most AI funnels accept anything with a sponsor and a demo, which guarantees a queue nobody can clear.

Four criteria do most of the work at intake. The value line: which cost, revenue, risk or capital line moves, and by roughly how much at the volumes the process actually handles. Reversibility: what happens if the system is wrong, and whether a bad action can be undone before it reaches a customer, a ledger or a regulator. Data and control feasibility: whether the inputs exist at usable quality and whether the decision can be made inside an existing permission and audit boundary. Owner availability: the named business person accountable for outcomes and exceptions for at least the next eighteen months, rather than the sponsor who wants the announcement.

Then set a takt time, the heartbeat of the line. Ninety days from intake to a first production decision is a defensible default, and the mid-market evidence suggests it is achievable. The number matters less than having one, because a factory with no cadence accumulates work in progress instead of output.

The rule that changes behaviour most is the one that bans pilot purgatory: once an idea passes a gate, it either proceeds to the next stage with funding or it is retired. It does not return to pilot for another cycle. Organisations that allow unlimited extensions are not being prudent. They are deferring a decision and paying for the delay in platform capacity, security review and executive attention.

Fund stages, not projects

Project funding and AI delivery are badly matched. A project is funded once, against an estimate, and then defended by the people who wrote the estimate. A product is funded in stages, against evidence, and can be stopped cheaply at any of them.

Stage-gated funding changes the conversation at the investment committee. Each stage carries a small budget, a named owner and a defined exit. Discovery buys the value hypothesis and the control assessment. Prototype buys an evaluation set and a measured baseline. Limited production buys real users, telemetry and run-cost data on a deliberately narrow scope. Scale buys capacity, support and integration hardening. Nothing buys the next stage except the evidence the last stage was supposed to produce.

Gartner’s research draws a useful line here. The organisations it classifies as high performers — those that track ROI continuously, treat AI as a portfolio of value, and reallocate or discontinue underperforming initiatives — reported positive returns in 81% of their AI initiatives. Low performers could not determine the rate of return for 29% of theirs. The difference is not analytical sophistication. It is the willingness to stop things.

That willingness has a financial edge in 2026. Roughly one in five respondents to McKinsey’s state of AI survey said AI operating costs, including token spend, had already constrained how much AI their organisation uses. A stage gate that cannot see run cost is not a gate. Gartner’s expectation that over 40% of agentic AI projects will be cancelled by the end of 2027, on escalating costs, unclear business value or inadequate risk controls, describes exactly the three failures a disciplined gate is designed to catch early.

Name the operator before you build

The most common failure I encounter has little to do with the model. It is a good pilot whose team dispersed.

Pilots are usually staffed by an innovation group, a data science team or an enthusiastic product manager with borrowed engineers. They end with a demonstration, a report and a decision to scale, at which point the borrowed people return to their day jobs. What remains is a service in production that nobody can explain, monitor, tune or take down.

The fix is procedural and unglamorous. No workload enters limited production without a named operator, an agreed support arrangement and a run-cost owner. The operator is a team rather than a person, and it appears in the same on-call rota as the rest of the estate. Its obligations include keeping the evaluation set current, watching drift and exception rates, owning the prompt and configuration versions, and holding the standing authority to reduce autonomy or switch the service off.

This is the point at which most AI portfolios shrink, and that is a feature rather than a problem. Work that cannot attract an operator is not a product. It is a demonstration with a budget.

Make evidence the currency of every gate

A product factory runs on evidence, and the evidence a gate requires should be boring, consistent and identical for every product.

On the delivery side, DORA’s 2026 ROI research is a useful corrective for teams expecting immediate returns from AI-assisted engineering. It describes AI as an amplifier of the system it lands in, with a J-curve in which measured output dips before it improves as teams absorb a verification tax on generated work. The same research reports productivity gains of 35–40% on simple, greenfield tasks and around 10% or less on complex legacy code. A CIO who budgets for the dip, and measures where the verification work lands, will get further than one who expects the first quarter to pay for itself.

Measure the factory, not the pilots

If you want to know whether you have a product factory, stop reporting the number of pilots and report the machine.

These are fitness functions for an operating discipline rather than for model quality. They also give a board something better than an activity count: a conversion rate it can hold management to, quarter after quarter.

The leadership test

Ask for the last ten ideas that entered your AI funnel, what happened to each, and what each one costs to run this month. If the answer is a list that is still running, or a set of numbers that would take a fortnight to assemble, you have a pilot factory wearing a product factory’s language.

The first decision is not a platform, a model or a centre of excellence. It is the rule that nothing enters the funnel without an owner, nothing leaves a stage without evidence, and nothing stays in the portfolio without a run cost someone can produce on demand. Factories are remembered for what they ship, and for what they stop.

Sources


Previous Post
You Contracted for a Capability, Not a Change Process
Next Post
The Agent Portfolio: Why Enterprises Need Fewer Agents with Clearer Owners