Skip to content

Every New Model Is a Production Change: Governing the AI Upgrade Cycle

Published: at 07:35 AMSuggest Changes

A product owner sees that a model provider has released a more capable option. The procurement team sees a lower price. Engineering sees a one-line configuration change. The customer-service operation sees neither of those things. It sees an assistant whose tone, refusal behaviour, retrieval use, tool choices and response time may change.

That gap is where avoidable AI incidents begin. A model replacement is often treated as an infrastructure upgrade because the application code still deploys and the API call still returns text. In a live service, it is a behaviour change. The service has acquired a different decision-making component, with different strengths, failure patterns and operating costs.

CIOs should govern model upgrades as production changes. That does not mean assembling a monthly approval board for every provider release. It means matching evidence, authority and rollout controls to the consequence of the service being changed. A low-risk internal drafting tool deserves a light process. A model that selects cases, drafts regulated advice, writes code or invokes enterprise tools deserves much stronger proof.

Start with the service, not the provider announcement

Model release notes are useful input. They are not release evidence for an enterprise application.

A provider may report better reasoning, lower latency or a new tool-use feature. Those claims describe tests chosen by the provider, under its own conditions. They cannot establish that a new model will apply a company’s current policy correctly, keep a customer workflow within its authority or produce code that fits an existing repository. The same model can improve one task and regress on another.

This matters because an AI service is more than a model. It includes system instructions, retrieval, tool schemas, policies, fallback logic, user interface, identity controls and the business workflow around it. A model upgrade changes that assembled service. It may produce different query formulations, call a tool in a different order, abstain less often, consume a different number of tokens or expose an assumption hidden in a prompt.

NIST’s Generative AI Profile frames risk management across the design, development, use and evaluation of AI products, services and systems. That is the appropriate unit for change control. The operational question is not whether the supplier’s new model is impressive. It is whether this service remains safe and useful for the people and process it affects. NIST AI 600-1

Create a model inventory before the first forced migration creates it for you. For each production service, record the provider, model identifier and whether it is a dated snapshot or an alias; the regions and endpoints used; dependent prompts, tools and retrieval assets; the business owner; the technical owner; the allowed action authority; and the provider’s retirement or support date. Include embedded AI in purchased software where it makes or informs material decisions.

The inventory turns a vendor notice into a manageable portfolio question. Which services are exposed? Who must decide? Which changes can be grouped because they use the same task contract? Which dependency will stop working if the team does nothing?

Classify the change by consequence

One upgrade policy for every model is lazy governance. It either delays harmless maintenance or permits risky changes with too little scrutiny.

Classify a proposed change against the service’s consequence, rather than against the marketing label of the model. A sensible classification considers who is affected, whether the system can act or only draft, the sensitivity of data, reversibility, regulatory obligations, transaction value, and the reliability of the existing fallback.

An internal writing assistant that cannot access confidential records or take action may need a smoke test, representative comparison and a named owner accepting the result. A knowledge assistant used by frontline staff needs freshness, grounding and escalation tests because confident-but-obsolete answers become an operational problem. A coding agent needs repository-specific checks, security tests, review evidence and staged use. A workflow that changes customer records, routes payments or triggers production actions needs formal comparative evaluation, explicit approval, limited rollout and a rehearsed reversal path.

The EU AI Act offers a useful discipline even where a service is outside its scope. Its obligations for general-purpose AI providers include technical documentation and instructions for downstream providers, while organisations deploying systems still need to understand the use they are putting into operation. A provider’s compliance material can support due diligence; it cannot replace the deployer’s responsibility for the workflow. European Commission

Make the classification visible in the change record. It should state the reason for the upgrade, services and users affected, expected benefits, risks that could worsen, evaluation required, approver, rollout scope, monitoring period and rollback trigger. This is deliberately ordinary change management. Ordinary is an advantage when the technology itself is moving quickly.

Compare the candidate against the service people use

The first gate is comparative evidence. Run the current configuration and the candidate through the same versioned task set. Keep prompts, policy rules, retrieval corpus, tool schemas, grader configuration and test environment stable unless one of them is part of the requested change. Otherwise, the team cannot tell what caused a difference.

The test set must include normal work and cases designed to expose harm: ambiguous customer requests, stale or conflicting sources, sensitive-data prompts, malformed tool responses, requests beyond the service’s authority, and cases where refusal or escalation is the correct result. For an agent, score the path as well as the final answer. Did it choose an authorised tool? Did it make an unnecessary retry? Did it preserve the evidence required by the reviewer?

Aggregate scores help leaders see a release at a glance. They must not hide a severe regression. A candidate that slightly improves average helpfulness while issuing unsafe tool calls in two high-impact cases has failed the release gate. Define non-negotiable thresholds for named policy breaches, prohibited actions, unsupported material claims and security controls. Record meaningful case-level changes for the people who own the business risk.

OpenAI’s guidance on evaluations argues for task-specific tests and continuous evaluation rather than relying on informal demonstrations. That principle is sound, though every organisation must set its own threshold and sample design. An evaluation set becomes valuable when it reflects the work the service is paid to do and the situations in which it must stop. OpenAI

Model upgrades also need a technical compatibility test. Check response structure, structured-output validity, tokenisation, context limits, rate limits, safety or refusal signals, tool-call formatting, SDK versions and observability fields. A response that is semantically better can still break downstream parsing or cause a cost spike. Anthropic’s migration guidance illustrates the point: a model transition can involve changes to thinking configuration, sampling parameters, message prefilling, tool versions and response content structure. These are application changes, not merely new names in a configuration file. Anthropic

Release progressively and keep a way back

Passing an offline suite earns a controlled rollout. It does not prove that live traffic, data distributions and human hand-offs will behave as expected.

Use a dated model version where the provider supports it. An unpinned alias can change outside the organisation’s release window, making diagnosis and rollback harder. If a provider only offers a moving alias, record that risk and increase production monitoring. Provider retirement schedules also matter. Anthropic’s published deprecation guidance, for example, advises customers to test replacement models well before a retirement date and distinguishes active, legacy, deprecated and retired lifecycle states. Build those dates into service ownership, rather than leaving them in a vendor newsletter. Anthropic

Begin with shadow comparisons where possible: run the candidate on live-shaped inputs without exposing its outputs or actions to users. Then release to a small, suitable cohort or a bounded class of work. Keep the existing model available during the observation period. For action-taking agents, start with recommendation mode or a tighter action budget before restoring full authority.

Monitor the service contract, not only infrastructure health. Track task-quality samples, escalation and override rates, grounding failures, policy blocks, refusal changes, tool-call errors, latency, token use, cost per completed task, customer correction and business outcomes. Segmentation matters. An average can hide a failure in a language, customer segment, product line or workflow step.

A rollback plan needs more than a previous model name. Confirm that the old version remains available, deployment configuration is versioned, caches and retrieval changes are understood, and downstream schema changes can be reversed. Set the people and thresholds that can halt the rollout. A service owner should not have to negotiate emergency authority while an unsupported answer is already reaching customers.

Google Cloud’s guidance for operating generative AI applications calls for versioning the chain, inputs, outputs and intermediate states, with monitoring for degradation and unexpected behaviour. That guidance is especially useful during a model change. Preserve the configuration and evidence that explain what the service did before and after the release. Google Cloud

Make model change a repeatable delivery capability

The durable answer is a small upgrade pipeline, not a heroic migration project repeated for each release.

Product teams should own the task contract, representative cases, acceptable failures and business outcome. Platform teams should make model inventory, version pinning, evaluation execution, telemetry, cost allocation and rollback mechanics easy to use. Security and risk functions should define the controls that cannot be traded away. Procurement and vendor management should surface retirement notices, contractual commitments and material service changes early enough for teams to act.

Measure the pipeline. Track time from notice to assessed impact, services with a current model inventory, percentage of material changes with comparative evidence, rollback success, post-release regressions, and production-discovered changes. These are fitness functions for an upgrade capability. They show whether the organisation is learning to change safely or merely becoming better at documenting surprises.

There is a temptation to wait until provider release velocity settles. It will not. Models, endpoints, safety behaviour, pricing and support dates will continue to move. The leadership test is therefore direct: when the next model notice arrives, can the accountable owner show what will change in their service, evidence that the candidate meets its task contract, who may approve the rollout and how customers will be protected if it regresses? If the answer is no, the organisation has outsourced a production dependency without taking responsibility for operating it.

Sources


Previous Post
The Agent Portfolio: Why Enterprises Need Fewer Agents with Clearer Owners
Next Post
Your AI Is Only as Current as Its Context: Managing Enterprise Knowledge for Agentic Work