Skip to content

You Contracted for a Capability, Not a Change Process

Published: at 12:50 AMSuggest Changes

A bank signs a three-year agreement for an AI assistant. The bake-off was rigorous: the model scored 91% on the bank’s own evaluation set, latency met the threshold, and legal cleared the terms. Eighteen months later the vendor has upgraded the model twice, retired the version that was tested, changed the price per token, and moved a slice of inference to a subprocessor the bank never reviewed. The capability named on the invoice is unchanged. The system in production is a different animal.

That gap is the defining commercial risk of enterprise AI, and it is a procurement failure rather than a technology one. When we buy AI, we tell ourselves we are buying a capability: a snapshot, a benchmark, a price. What we actually sign is a relationship with a vendor whose model, pricing and terms will move underneath us, at a cadence our contracts were never written to govern.

The vendors are not being dishonest about this. It is the nature of the product. A model is not a released artefact; it is a running service that improves, drifts and gets retrained. The mistake sits on the buyer’s side: treating a moving system as if it were static, and signing terms that assume the thing tested at selection is the thing running at renewal.

The change clause is the first thing to write

Most AI contracts contain no meaningful change clause. The vendor reserves the right to “improve” the service, and the buyer hears about material changes only if the vendor chooses to mention them. By the time an upgrade reaches production, the buyer has lost the ability to decide.

A government buyer has now shown what the alternative looks like. The US Department of Energy’s Acquisition Letter AL 2026-05, issued on 8 May 2026, requires contractors to give “advance written notice of material changes to model architecture, training data sources… safety policies, or content-moderation configurations” that are likely to affect the factuality, neutrality or safety of outputs. It is a narrow use case with a political motive, but the structure is exactly what any enterprise needs: the supplier must tell the customer, in advance, when something material changes. [US Department of Energy, May 2026]

Three terms belong in every AI agreement. First, an immutable model identifier pinned in production, so the buyer knows precisely which version is answering. Second, a notice period that exceeds the buyer’s tested migration time — if it takes eight weeks to re-evaluate and cut over, a two-week notice is not a notice, it is a fait accompli. Third, an overlap window in which the old version keeps running while the new one is assessed, with the right to stay pinned if the new version fails the buyer’s evaluation.

The clause also has to define “material”. Continuous tuning is the norm and should not trigger a conversation every fortnight. What matters is a change that could alter outputs the buyer relies on: a re-trained base model, a shift in safety or refusal behaviour, a change in how the system routes or caches requests, or a move to a new subprocessor. Ask the vendor to draw that line in the contract, not in a product changelog. Gartner’s January 2026 research on AI contracting is blunt about the consequence of getting this wrong: buyers must “strengthen contractual protection and create contingency plans for when vendors falter”, because pricing models and use cases are moving faster than the paperwork around them. [Gartner, January 2026]

Your data-use boundary is a contract term, not a toggle

Every AI vendor says it does not train on your data. The statement is worth exactly as much as the clause behind it, and the clause is often weaker than the sales conversation implies.

The contract has to separate four things vendors like to blur: your inputs, the outputs, the derivatives created from your inputs (fine-tunes, embeddings, indexes, caches), and the vendor’s underlying technology. Model training is a different activity from service improvement, and telemetry that keeps a service secure and reliable is different again. Each needs its own boundary. A vendor that will not distinguish them has either not thought about your risk or decided not to.

The subprocessor question is where this gets uncomfortable. DataGrail’s Privacy and AI Trends Report 2026, released on 27 May 2026, reviewed 2,400 business software providers that advertise AI capabilities and found that 63.6% did not disclose a third-party AI subprocessor in their legal documentation. Nearly a third of the AI systems that disclosed AI capabilities also reported at least one high-risk activity, such as sensitive-data processing or automated decision-making. The data processing agreement you signed may not describe the AI that is already touching your data. [DataGrail, May 2026]

The IAPP made the operational point in August 2026: when answers about data use differ between the sales presentation, the technical documentation and the contract, that is not a drafting quibble, it is a governance and risk problem, because the buyer is authorising something it does not fully understand. Fix it by requiring the vendor to name every subprocessor in the agreement, to notify before adding one, and to give the buyer a right to object. [IAPP, August 2026]

Buy the evidence, not the demo

A leaderboard score tells you almost nothing about whether a model will work in your workflow. Ask for evaluation evidence against your own task, your own data and your own failure modes — and ask for it in a form you can re-run.

The regulatory floor now exists, which makes this easier to demand. General-purpose AI model obligations under the EU AI Act became applicable on 2 August 2025, requiring providers to maintain technical documentation, publish a summary of training data, and pass information downstream to the organisations that build on their models. A vendor selling into Europe is already obliged to hold this material. Ask to see the parts that bear on your use. [European Commission, August 2025]

Two contractual rights follow. The first is a warranty that the model meets the agreed evaluation thresholds at the point of delivery, not merely at the point of selection. The second is the right to re-test after any material change, with the results shared. An evaluation you ran once at procurement is a historical document. An evaluation you can re-run on demand is a control.

Service levels for systems that do not behave like services

Traditional SLAs measure availability: is the service up, is it fast. Those metrics still matter, but they are the easy half. For a non-deterministic system the harder question is whether the outputs remain usable.

Define the service in terms your business can feel: the proportion of tasks completed without human rescue, the rate at which outputs fall outside agreed quality bands, drift against a golden set, tool-call reliability, and escalation latency. Set thresholds and attach remedies that change behaviour — service credits, the right to pin to a known-good version, the right to roll back, and ultimately the right to exit. A vendor that will only guarantee uptime is telling you where it has stopped taking responsibility.

Allocate liability in proportion to control

The vendor chooses the model, the training data, the safety tuning and the routing. When the system fabricates a fact, reproduces copyrighted material, or leaks data through a prompt-injection flaw, the party that controls those decisions should carry the risk. Too many enterprise agreements do the opposite: a broad disclaimer of responsibility for outputs, paired with a liability cap small enough to be meaningless, signed by a buyer who was told the tool was fit for their workflow.

Push back on blanket output disclaimers. Accept that a vendor cannot warrant that a probabilistic system will never be wrong; do not accept that it bears no consequence when it is wrong in a predictable, AI-specific way. The test is simple: where the vendor controls the risk, the vendor should share it.

Exit is a design requirement

AI lock-in is sharper than ordinary software lock-in because what you cannot take with you is often the thing you built. Prompt libraries, evaluation sets, retrieval indexes, embeddings and fine-tuned behaviour are not “customer data” in the traditional sense, and most contracts do not name them. If they are not named, they belong to the vendor by default.

A workable exit clause covers four categories: your raw data; the artefacts your teams created; the outputs you generated, with rights that outlive the relationship; and the vendor’s deletion obligation, ideally with written certification. It specifies what stands in for the model when the model cannot move — training and fine-tuning datasets, system instructions, performance logs — so a replacement vendor has something to work with. It fixes exit-services pricing at signing, while neither party has leverage to distort it, and it lists the triggers that activate the plan: contract expiry, a decision to switch or insource, model deprecation, a change of control, a pricing change outside agreed bands, and a security incident.

The decision before you sign

Most of this is knowable before signature. The vendor either has a change-notice clause that exceeds your migration time or it does not. It either names its subprocessors or it will not. It either has an exit plan with portable artefacts and fixed pricing or it is improvising.

The leadership test is to ask for four things in writing before the contract goes to signature: the change-notice clause, the data-use and subprocessor schedule, the last two evaluation reports for your use case, and the exit plan. If the vendor cannot produce a change notice longer than your tested migration time, you have not bought a capability. You have bought a subscription to someone else’s roadmap, priced as if it were yours.

The decision is then clean. Either make the change clause and the exit plan conditions of signing, or price the risk into the deal and record that you chose to carry it. What you should not do is sign a static contract for a moving system and discover the difference eighteen months later, when the version you tested no longer exists.

Sources


Previous Post
The Compute Bill Comes Due: Why AI Roadmaps Now Have a Capacity Ceiling
Next Post
The Agent Portfolio: Why Enterprises Need Fewer Agents with Clearer Owners