The CIO has two teams waiting. One needs a model endpoint and proof its customer-service assistant is controlled. The other wants to test whether retrieval, a workflow, or a tool-using agent will solve an operations problem. Security wants identity, data boundaries and logs. Finance wants to know who is spending what. Architecture proposes an enterprise AI platform.
That proposal can either remove weeks of avoidable delivery work or freeze the two teams behind a new internal product they did not ask for. The difference is leadership discipline. A platform is justified when a capability recurs across live products and the absence of a shared answer creates material risk, cost or delay. It is a poor substitute for product teams still learning what they are building.
I have seen this pattern with integration hubs and developer portals. Central teams spot duplication, draw a target architecture and centralise decisions that were still product hypotheses. The catalogue grows. Adoption has to be negotiated. Exceptions multiply. Meanwhile, identity, secrets and audit evidence remain uneven.
Enterprise AI will repeat that mistake unless leaders separate delivery foundations from application judgement.
Begin with the constraint that is hurting delivery
Several AI pilots do not make a platform case. The evidence is in the delivery path. Follow one AI-enabled service from a business request through data access, model selection, evaluation, security review, deployment, monitoring and incident response. Identify controls rebuilt by each team and queues that protect the business.
A repeated problem is easy to recognise. Each team is negotiating access to the same model providers. API keys are created in ad hoc ways. Engineers cannot tell which product generated an invoice spike. Prompt, tool-call and model-version records are absent when an incident occurs. A data classification decision is made three times, by three people, with three different results. Those are platform concerns because local variation carries little business benefit and creates work that no product team should have to repeat.
Other questions must stay close to the workload. Should a claims assistant retrieve policy documents or ask a human reviewer? Which cases belong in its evaluation set? What constitutes an acceptable failure? When should a procurement agent stop and escalate? The answer changes with the process, the data, the people affected and the cost of getting it wrong. A central team can provide expertise and templates. It should not take away the product owner’s decision right.
AWS separates reliable infrastructure, foundation-model selection, security and governance, and repeatable application patterns in its guidance for enterprise-ready generative AI. That separation is useful because it exposes distinct responsibilities. It does not mean each layer belongs in a single, large internal product. AWS
The first executive decision is simple: name the recurring constraint, teams affected and evidence that a shared capability will improve their route to production. If nobody can do that, do not fund a platform programme yet.
Build a paved road that teams can actually use
A first platform increment should be small enough to prove its value. Its job is to make the governed route easier than an improvised one.
That generally means a shared set of foundations: workload and user identity; approved model endpoints with rate limits and usage records; managed secrets; network and data-classification controls; telemetry for requests, latency, failures, tool calls and spend; reusable deployment templates; and a baseline pattern for evaluation and release evidence.
These controls matter because they establish a common operational language. When a service changes model, when a tool call behaves unexpectedly, or when a regulator asks how a decision was supported, teams should be able to find the owner, version, evidence and rollback path without assembling it from ticket comments.
Microsoft’s landing-zone guidance draws a sensible boundary between shared resources such as networking, identity, policy and monitoring, and the application landing zone where the workload operates. It recommends that the workload own its AI resource rather than defaulting to a central resource. The product names are Azure-specific. The operating lesson is broader: central teams supply foundations; product teams remain accountable for the service they run. Microsoft
This has a consequence for the platform team’s service design. A team that needs an approved model should receive it through workload identity and a documented self-service route, not a ticket queue. A team that needs a different model should face a proportionate review with clear decision criteria, rather than an informal exception process. If every useful change requires a meeting with the platform team, the organisation has created a dependency, not a paved road.
Keep product learning where the work happens
The most expensive centralisation mistake is forcing every team into the same retrieval pattern, agent framework or prompt template. These choices look standard from a distance. They are not standard in operation.
Take a knowledge service. A central index may appear efficient until teams ask who owns the documents, which version is authoritative, how fast an update must appear, which users may see it and how sensitive material is excluded. A policy assistant, a developer assistant and an operations triage tool can all use retrieval. They do not share the same corpus, access rules, freshness requirements or failure consequences. A central index that ignores those differences creates an attractive demonstration and a hard audit problem.
The same applies to agent runtime. AWS’s guidance on multi-tenant agentic architectures identifies authentication, discovery, tenant isolation, resource management, data ownership, monitoring and testing as shared concerns as agents spread. It also recognises that deployment models vary. That is the right premise: standardise the capabilities that make an agent controllable; do not assume every agent belongs in one centrally owned runtime. AWS
A useful test is whether a proposed platform capability changes the outcome for more than one product team today. A common model-access gateway often passes. A central agent orchestration service designed for forecast use cases usually does not. Keep the latter with the team that needs it until repeated operational evidence makes the pattern clear.
This approach avoids two failures. One is the grand platform built in anticipation of hundreds of use cases, complete with a service catalogue and little pull from delivery teams. The other is free-for-all duplication, where every team chooses a gateway, logging scheme, vendor contract and confidential-data policy. The first wastes central investment. The second creates control debt that surfaces at the worst possible moment.
Work with two or three live teams that have materially different needs. Give them a shared access and observability slice. Watch where the route works, where it does not, and which exceptions recur. Build the next capability from that evidence.
Shared evidence, local accountability
Model choice is where platform ambition often overreaches. A shared service can provide approved models, routing controls, version records, quota management and consistent measures of latency and cost. It cannot decide what good looks like for every product.
A contract summariser, a developer assistant and an operations triage tool fail differently. One may omit a clause, another may introduce a defect, and the third may send a case to the wrong queue. Each needs representative test cases, defined escalation routes and an acceptable-error threshold set by the person accountable for the process.
NIST’s Generative AI Profile places trustworthiness across the design, development, use and evaluation of AI products, services and systems. That should prevent governance becoming a central approval ceremony. The platform sets the minimum evidence that a release must carry. The workload team demonstrates that its service meets that threshold for its own decisions and users. NIST
Google Cloud makes a similar point in its guidance for operating generative AI applications. Data curation, prompt engineering, model tuning and grounding sit in the application lifecycle alongside adapted DevOps and MLOps practices. A platform can provide the tools and standards. Feedback from users, product quality and domain-specific failure analysis remain delivery work. Google Cloud
The release record should make the split visible. Every production deployment needs a named business and technical owner, model and version, evaluation evidence, data boundary, intended action authority, monitoring signals, and rollback path. The platform should make that record easy to produce and hard to omit. Product teams should own its truthfulness.
Fund increments against measurable fitness functions
Platform adoption is a weak measure on its own. Teams sometimes adopt a central service because policy leaves them no choice. Measure whether it improves delivery and control.
Track the time from a team’s request to a governed production deployment. Track the proportion of workloads using standard identity, telemetry and cost allocation without bespoke engineering. Measure policy exceptions and their cycle time. Test whether a team can change a model or roll back a release without opening a platform ticket. Compare unit cost and operational reliability with the route the team used before.
Then measure the cost of centralisation with equal honesty. How long does a new model, connector or configuration take? How many workarounds are appearing? Is the platform team carrying an on-call burden it can support? Does the service catalogue expand while product teams wait longer? Those signals tell a CIO when to stop adding features and repair the service.
Decision rights need to be explicit. The platform team owns the paved road, service levels, common controls and roadmap. Product teams own business outcomes, evaluation, data use and production behaviour. A lightweight forum can review requests for a new shared capability when two or more teams bring evidence of the same constraint. It should not become a permission gate for early learning.
For most organisations, the next move is straightforward: give two product teams governed model access through workload identities, record usage and cost by product, enforce existing data and network controls, and capture a release record with model version, evaluation evidence, owner and rollback path. Leave application logic, retrieval design and agent orchestration with the teams close to the work.
Run that arrangement long enough to see whether it reduces delivery time and improves control evidence. Expand only where the proof is there.
The leadership test is simple. When a product team can ship an AI service safely, explain its behaviour, account for its spend and reverse a bad release without begging a central team for help, the delivery architecture is doing its job. If the platform makes those things harder, it is a new bottleneck with better branding.
Sources
- Amazon Web Services, Building an enterprise-ready generative AI platform on AWS, June 2025.
- Amazon Web Services, Building multi-tenant architectures for agentic AI on AWS, July 2025.
- Microsoft, Baseline Microsoft Foundry chat reference architecture in an Azure landing zone, accessed before the scheduled publication date.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 26 July 2024.
- Google Cloud, Deploy and operate generative AI applications, last reviewed 19 November 2024.