A chief audit executive I worked with signed a clean opinion on an AI-assisted credit decisioning control in March. By September, the control no longer existed in any meaningful sense. The vendor had upgraded the underlying model, a product team had rewritten the prompt that shaped the decision summary, and the mix of data feeding the retrieval layer had drifted. None of it raised a flag. The evidence pack was immaculate, and it described a system that had already stopped running.
That is the trap the third line is walking into across the enterprise. Internal audit was built to assure controls that behave the same way on the day of testing as they did on the day of design. Sample a population, test whether the control operated, extrapolate to a conclusion, repeat next year. The method assumes a stable control and a stable population. AI systems offer neither.
The position worth defending is uncomfortable for audit committees. An annual opinion on an AI control is increasingly a statement about a configuration, not a control. It tells the board that a system passed on one afternoon and says nothing about the eleven months that follow. Treated as assurance, it is a liability dressed as comfort.
The control changes while you are testing it
Deterministic controls fail in legible ways. An approval threshold is set at the wrong level, a segregation-of-duties rule is misconfigured, a reconciliation breaks when an upstream field is renamed. The control is either designed correctly and operating, or it is not. Testing is a question of evidence.
AI systems behave differently. Their output is a function of a model version, a prompt, the retrieved context, the tool permissions in force and the data distribution at the moment of the call. Change any one of those and the effective control changes with it. The same input can produce a correct answer on Monday and a plausible but wrong one on Friday, with no code change and no ticket. The population is not stable, so a sample drawn from it is not a sample in any statistical sense the audit profession recognises.
The consequence is that the classic audit assertion — the control operated effectively throughout the period — cannot be made from a point-in-time test of an AI system. You can test a snapshot. You cannot sample your way to an opinion about a system whose behaviour is a moving target.
Continuous evidence replaces the pre-audit pack
The first demand to make of management is a change of artefact. Stop accepting the evidence pack assembled in the two weeks before fieldwork. Ask instead for evidence that is produced continuously as the system runs: drift and performance dashboards, evaluation runs with their scores, override and exception logs, incident records, and the alert history for threshold breaches.
The Institute of Internal Auditors already places internal audit in exactly this seat. Its Artificial Intelligence Auditing Framework, updated in September 2024, sets out internal audit’s role as the third line within the Three Lines Model, providing independent assurance over how AI is governed, managed and monitored, alongside an advisory role. [IIA, September 2024] The framework is explicit that assurance covers the processes management has established to run AI, not merely the outputs it produces.
The practical shift is from inspecting a document to interrogating a pipeline. Does the monitoring exist, is it trustworthy, does it cover the failure modes that matter, and did anyone act on it? A drift alert that nobody actioned is a finding about the control environment, whatever the model’s accuracy score says.
Assure the control system, not the output
The second demand is a change of target. Auditing individual AI outputs is a losing game: there are too many, they vary, and a clean sample proves little. The defensible target is the system of controls that keeps outputs inside tolerance — and that system looks a great deal like the machinery regulators already expect for models.
The Prudential Regulation Authority’s supervisory statement on model risk management, SS1/23, published in May 2023 and effective from May 2024, sets out five principles: a model inventory and risk-based tiering; governance with board-approved risk appetite and an accountable senior manager; documented model development and use; independent validation that provides ongoing challenge; and defined mitigants, including restrictions on a model’s use when deficiencies are found. It explicitly extends to risks from machine learning used in models, and expects reporting on the effectiveness of model risk management to reach the audit committee. [PRA, May 2023]
Read those principles as an audit programme rather than a banking rulebook. The questions transfer directly to any enterprise: is there an inventory, is it tiered by consequence, who can change the system, who validates it independently, and what happens when it underperforms?
The Committee of Sponsoring Organizations of the Treadway Commission made the same move for generative AI. Its February 2026 publication, Achieving Effective Internal Control Over Generative AI, adapts the established internal control framework to a technology it describes as bringing opaque reasoning, model drift and frequent configuration changes. Its roadmap runs govern, inventory, assess, design, implement and monitor, with templates for risk assessment and control testing. [COSO, February 2026]
The logic for the third line is straightforward. If the control system is sound — changes are governed, validation is independent, monitoring is live and breaches have consequences — then individual output errors are tolerable and self-correcting. If the control system is weak, a clean sample is luck dressed as assurance.
Sampling behaviour that refuses to hold still
Sampling is not dead, but its purpose changes. Instead of estimating a population error rate, audit uses sampling to probe how a non-deterministic system behaves across conditions and how it fails at the edges.
Three techniques earn their place. First, stratified sampling across input classes: the routine cases, the rare cases, and the cases the system was never designed to see. Second, repeat sampling of identical inputs to measure variance — if the same prompt yields materially different answers, the control has a stability problem that a single test would miss. Third, adversarial sampling, where auditors construct inputs designed to break the control rather than confirm it, in the spirit of red-teaming the guardrails.
Add a fourth: sample the override log. The moments when a human overrode the system, approved an exception or corrected an output are where the real control boundary sits. A control that depends on constant human rescue is not operating; it is being operated.
Report the result as a range with stated failure modes, rather than a pass rate. “In 3% of adversarial cases the system produced an unverified recommendation that would have been actioned without review” is an opinion a board can use. “No exceptions in a sample of 40” is not.
The register that makes any opinion possible
If the audit committee funds one new artefact this year, it should be the change register. Internal audit should require a single, complete log of every material change to a model, a prompt, a retrieval source, a tool permission or a guardrail — each entry carrying a date, an owner, the reason, the test evidence and the rollback plan.
This is where the PRA’s expectations and the COSO roadmap converge. Model changes are meant to be documented, validated and subject to performance monitoring; configuration changes in generative systems are precisely what erodes control over time. [PRA, May 2023] [COSO, February 2026]
Without the register, no audit opinion on an AI control can hold, because nobody can answer the only question that matters across a period: what changed since the last time we looked? With it, audit gains a spine. It tests the register itself — is it complete, are high-risk changes pre-approved, are post-change evaluations run, is independent re-validation triggered when the change is material?
What the board should actually receive
Boards now carry a heavier, more explicit obligation. The Financial Reporting Council’s UK Corporate Governance Code 2024, published in January 2024, requires through Provision 29 that the board declare the effectiveness of its material controls as at the balance sheet date, covering financial, operational, reporting and compliance controls, for financial years beginning on or after 1 January 2026. [FRC, January 2024] A declaration of that kind cannot be supported by a stale AI control opinion.
The appetite for this work is real, and so is the gap. In the Deloitte and Center for Audit Quality Audit Committee Practices Report published in September 2026, 75% of audit committee members ranked AI governance among their top three priorities, up from 35% a year earlier, yet only 56% were confident in their committee’s ability to oversee it. [Deloitte and CAQ, September 2026] The shortfall is an assurance gap, and committees cannot close it with a better dashboard.
A board report on AI should therefore drop the finding count and carry four things: residual risk by tier, the age of the newest evidence behind each material control, the material change events since the last report, and the specific decisions requested. The most telling number on the page is the second one. A control whose freshest evidence is six months old is not controlled; it is remembered.
The decision for this year
The decision in front of the audit committee is how to scope AI assurance for the coming year, and the honest answer is that it cannot be scoped as a once-a-year engagement.
One test settles it. Ask the chief audit executive to produce, for the three AI systems with the greatest customer or financial consequence: the date each was last independently validated; every material change since; the continuous monitoring evidence for the last quarter; and a current residual risk statement for each. If any of those four cannot be produced on demand, the committee should decline to accept an annual AI control opinion, fund continuous assurance instead, and make the change register a standing precondition for any AI system that carries real consequence.
Assurance over a system that changes itself is a discipline of watching, not a ritual of visiting. Committees that accept the ritual will eventually discover the difference at the worst possible moment.