The operations director asks an internal agent why a customer exception was declined. The answer is clear, polite and wrong. It cites a policy that was replaced six weeks ago, because the old PDF sat in a shared drive beside the approved version. Nobody deliberately gave the agent bad information. Nobody owned the retirement of the old document either.
That is the enterprise knowledge problem in its most expensive form. Once an AI assistant only drafts text, stale knowledge is awkward. Once an agent can advise staff, route work, prepare a decision or call a tool, stale knowledge becomes an operating risk. The model may be capable. The workflow may have sensible controls. Yet the service is still working from an unreliable picture of the business.
CIOs should stop treating this as a retrieval tuning exercise. It is a management problem: decide what information an agent may rely on, who keeps it current, how authority is resolved when sources disagree, who may see it, and how a reviewer can trace a consequential answer back to its evidence.
An enterprise does not have one knowledge base
Most organisations have many overlapping versions of the same truth: policies approved by legal, procedures maintained by operations, frontline workarounds, product notes in a wiki, customer commitments in contracts, and decisions preserved only in an inbox. Search makes this material easier to find. It does not make it authoritative.
Agents treat retrieved material as working context. Anthropic describes context engineering as curating information within a finite attention budget. Its guidance favours the smallest high-signal set for the outcome required, rather than filling a context window with every document that might be relevant. Anthropic
More documents do not make an agent better informed. They can make it less dependable by mixing current policy with obsolete material, local practice with corporate standard, or a draft contract with the signed version.
I have seen this in transformation programmes. A team would publish a process standard, then leave the old process map and a dozen local variants in the same repository. The people who knew the difference made reasonable choices. A new employee struggled. An agent has the same problem faster.
A policy assistant needs one answer: which source wins? A product-support agent must distinguish a live service notice, release note and archived troubleshooting guide. A procurement agent needs the executed agreement and current delegation rules, not the slide deck that proposed them. The source hierarchy belongs in the service design, not in the model’s guesswork.
Treat knowledge as a product with an owner
Every domain used by an agent needs a named knowledge owner. That is not necessarily the person who uploads files. It is the leader accountable for the truth, approval path and review cadence of the material the service is allowed to use.
For an HR policy assistant, HR owns policy authority while legal may approve specific clauses. For a technology runbook agent, the service owner owns the operational procedure, while security owns mandatory controls. For a customer-support assistant, product and operations need an explicit agreement on whether a product release note, service status update or contractual commitment takes precedence.
The owner should maintain a modest but disciplined record for each approved knowledge domain:
- its business purpose and the decisions or tasks it supports;
- the authoritative systems and document types;
- a source owner and a review or expiry date;
- intended users and access classification;
- freshness expectations, including how quickly a material change must appear;
- conflict rules when two sources say different things; and
- the conditions under which the agent must abstain or escalate.
This is not a call to catalogue every page in the company. Start with the information needed by a live workflow. A finance agent preparing an access-review pack does not need the whole intranet. It needs approved entitlement records, the relevant policy, current role definitions and a route to flag missing evidence.
The hard part is deciding what to exclude. Drafts, personal notes, expired procedures, unapproved presentations and copies downloaded to local team drives may be useful to people. They are poor source material for an agent that creates a business record. If a source is not fit for a colleague to rely on without additional interpretation, it is not fit for autonomous retrieval.
Design provenance into the answer and the action
A user should be able to see why an agent gave a material answer. A reviewer should be able to reproduce that answer against the relevant version of the source. Operations should be able to find which knowledge item shaped an action when something goes wrong.
That requires more than a link at the bottom of a chat response. The service needs to retain the source identifier, version or effective date, retrieval time and the policy or workflow rule that allowed the response or action. Where an answer combines several sources, it should show the important ones rather than create a false impression of certainty.
NIST’s Generative AI Profile recommends documentation and retention of test, evaluation, validation and verification history, and calls out data quality, integrity and provenance as matters for measurement and monitoring. That is useful guidance for knowledge-backed agents. The question is not simply whether the model produced plausible words. It is whether the organisation can show the evidence and controls behind the result. NIST
Consider a service-desk agent that proposes a resolution for an employee. A good design tells the employee which approved procedure it used and when that procedure was last reviewed. A stronger design detects that the procedure has expired, declines to invent an answer and routes the case to the service owner. The visible answer may feel less magical. It is far safer to operate.
Provenance also changes content maintenance. When an outdated article causes a correction or escalation, the team should identify the owner, correct the item, retire the old version and add the failure to the evaluation set. That closes the loop between knowledge governance and product quality.
Access boundaries travel with the knowledge
Knowledge quality is only half the problem. An agent that finds the right answer in the wrong repository can expose sensitive material or influence a decision with information the user was never entitled to see.
Many retrieval demonstrations begin with a broad corporate corpus because it is convenient. Production design needs the opposite starting point. Define the user, agent and workflow identity. Then retrieve only from sources each is authorised to access. The fact that a document is useful to an executive, lawyer or security team does not make it suitable context for a general employee assistant.
OpenAI’s 2025 agent tooling announcement illustrates the operational pattern: file search supports metadata filtering, and its example uses separate vector stores for different user groups so answers can reflect account settings and user roles. That is vendor material, not a universal architecture. The principle holds across platforms: access policy must constrain retrieval before the model sees the content. OpenAI
Keep the controls comprehensible. Product teams need to know whether their service retrieves from a copy of source material, a live system, or a curated index; how permissions are enforced; whether prompts and retrieved passages are retained; and who can investigate a disputed answer. Security teams need to test those controls against indirect prompt injection and accidental disclosure. Domain owners need a simple way to remove or correct a source.
Do not use a central vector store as a shortcut around these questions. It can reduce infrastructure duplication, but it cannot erase separate data owners, retention periods, confidentiality labels or legal duties. Centralise the capabilities that make retrieval safe and observable. Keep authority for content and business risk with the domain that owns the work.
Freshness needs a service level, not an annual clean-up
A knowledge base decays at different speeds. Product documentation may change with every release. Operational procedures change after incidents. HR policies have a formal review cycle but may require immediate updates after a regulatory change. A long-lived engineering principle can remain useful for years.
Give each domain a freshness expectation that reflects the consequence of error. A customer-facing returns policy may need same-day propagation after a change. A low-risk internal explainer may tolerate a quarterly review. The accountable owner should be alerted before the expiry date, and an item without a valid owner or review date should become ineligible for high-impact agent workflows.
That does not mean every document needs manual reapproval. Use event-driven updates where the authoritative system already records publication, effective date and retirement. Use scheduled checks where it does not. Measure the lag between a material change and its availability to the agent. Track answers corrected because of superseded knowledge, source conflicts found in review, abandoned content, and cases where the agent correctly refused to answer.
The refusal metric matters. An agent that admits it cannot find an approved current source has protected the organisation. A service that produces a confident synthesis from stale material has hidden the control failure.
Start with one bounded knowledge contract
The first move should be small and operational. Pick a workflow with a clear owner, a finite source set, meaningful but reversible value and a reachable group of reviewers. An internal policy assistant for a defined employee process, an access-review pack generator or a support tool for a narrow product line are reasonable candidates.
Before building the interface, agree the knowledge contract. Name the source systems, authority order, classification rules, freshness target, escalation cases, source evidence shown to the user, and the person who can withdraw the service if those conditions fail. Create a small evaluation set containing current questions, deliberately stale material, conflicting documents, unauthorised requests and questions with no approved answer.
Run it with real reviewers. Watch where people correct the source selection, where retrieval misses a current rule, and where a business term means something different across teams. Repair the knowledge process before expanding the agent’s authority. The work will expose content debt that has existed for years. That is a useful result, not a reason to hide it behind a better prompt.
The leadership test is straightforward. When an agent gives advice or takes a step in a business process, can the accountable owner show which approved source informed it, why the user was permitted to see it, when it was last valid and who will correct it tomorrow? If those answers are vague, the organisation has built an articulate search box. It has not built a dependable agent.
Sources
- Anthropic, Effective context engineering for AI agents, 29 September 2025.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 26 July 2024.
- OpenAI, New tools for building agents, 11 March 2025.