AI Governance Explained: Accountability, Risk, and Evidence
A test score leaves a decision unanswered
An AI assistant drafts replies about a shop’s refund policy. The team reports that 190 of 200 test replies met its answer-quality rubric. Can the assistant now send replies without review, or issue refunds? Neither permission follows from that score. The tests may cover drafting while the proposed release adds actions, different customers, and consequences that were never evaluated. These numbers and the shop throughout this article are fictional teaching examples.
AI governance establishes how an organization decides which AI uses it will permit, who is answerable for those decisions, what evidence is required, and how decisions change when conditions change. It covers the working system: data, model, application, people, and process. Technical evaluation supplies observations about behavior. Governance connects those observations to an authorized decision about a particular use. No model mathematics or programming is required here.
Define the use before judging the model
Give the proposed use a concrete boundary. For the shop, the assistant drafts policy explanations for trained support staff using approved documents. Staff check each draft before sending it. The assistant cannot change orders or initiate payments. State the supported languages, source documents, intended users, and cases routed to a person. “Customer-service AI” leaves too much unstated to evaluate or approve.
Compare this proposal with a feasible alternative, such as staff searching the same policy manually. The intended benefit might be shorter handling time without more incorrect advice or unresolved complaints. Measure the whole workflow, including checking and correcting drafts. A faster model response is not evidence that staff finish cases faster. If the benefit disappears once necessary review is included, changing or declining the use is a valid decision.
Ask who experiences the result, not only who buys the tool. A customer may miss an available refund after receiving incorrect advice; staff may face a growing queue of difficult escalations. Include those perspectives when deciding what failures matter. A privacy review examines permitted data handling, and a security review examines exposure and misuse. AI governance must also decide whether the resulting service is appropriate for its intended use.
Keep an inventory of AI uses so these boundaries remain findable. A useful entry connects the service owner, purpose, affected people, model or supplier, data sources, allowed actions, deployment status, and latest decision record. Two applications using the same model can require different decisions. The inventory makes it possible to find which deployed uses need review when a shared component changes.
Make risk a scenario someone can investigate
A useful risk statement connects a condition, a failure, and a consequence. “AI can be wrong” is too broad to assign work. “An outdated refund paragraph is retrieved, the assistant presents it as current, and a customer receives incorrect eligibility advice” points to source maintenance, retrieval checks, and review. Risk assessment considers how likely the scenario is and how serious its consequences could be, including uncertainty in both judgments.
| Scenario in this proposed use | Control to investigate | Evidence to seek |
|---|---|---|
| Obsolete policy reaches a draft | Remove obsolete sources and test refresh paths | Versioned retrieval tests after a policy change |
| An unsupported language produces misleading advice | Route unsupported cases to qualified staff | Routing tests and review of missed cases |
| Reviewers send incorrect drafts under time pressure | Provide sources, sufficient review time, and escalation authority | Observed reviewer performance under realistic workload |
A control is a measure intended to reduce risk. Its presence in a design document does not show that it works. Residual risk is the risk remaining after the controls actually applied. For example, fresh source documents do not ensure that the assistant interprets an exception correctly. Record what remains unresolved and who may decide whether that remaining exposure is acceptable for the defined use.
A simple high/medium/low scale can help prioritize work if the team defines it consistently. It does not turn an uncertain judgment into a measured probability. Nor should a good average score cancel an unresolved severe failure. Agree on decision criteria before examining release results, including failures that block release and missing evidence that requires more work.
Risk-based decisions compare alternatives using possible harms, benefits, and uncertainty. Avoiding a use, narrowing its permissions, improving controls, or accepting a documented residual risk are different responses. A numerical expected loss can inform that comparison when probabilities and consequences have a defensible basis. It cannot represent every harm or replace a release-blocking safety condition. Outsourcing work or buying insurance may allocate some costs, but does not remove the harm to an affected person or the organization’s decision responsibilities.
High-stakes uses can materially affect health, safety, rights, or access to important opportunities. The stakes depend on the actual decision, its scale, and how easily an error can be corrected. A hiring rejection and a clinician’s treatment recommendation call for different evidence from a searchable internal note. More consequential uses warrant stronger domain evaluation, scrutiny of affected groups, and a practical way to challenge or correct outcomes. Calling a whole industry “high risk” is too coarse to specify those controls.
Assign work and decision authority separately
Responsibility identifies who carries out work; accountability identifies who must answer for a decision and ensure follow-through. In this example, engineers run evaluations and repair retrieval, while an authorized service owner decides whether the supported workflow may launch within organizational policy. The owner must have authority and resources to impose conditions or stop the service. Naming someone without those powers leaves a gap.
Domain specialists define correct policy handling. Privacy and security specialists examine relevant controls. Operators need an escalation route and authority to apply the agreed stop conditions. A reviewer who can challenge the development team helps expose weak evidence; the degree of independence should match the consequences and organizational requirements. These functions need clear assignments, not necessarily a separate department for each one.
Buying a model does not assign these decisions to the vendor. The supplier can describe its component and available tests; the shop still needs to understand its own documents, permissions, staff behavior, and customers. Record what the supplier can change, which changes can be detected, and what the shop will do when required evidence is unavailable.
Read the evidence at the level of the claim
Return to the fictional result: 190 acceptable replies out of 200 is 95% on that test set under that rubric. Suppose 180 ordinary cases all pass, but only 10 of 20 exception cases pass. The overall result remains 95%, while exception-case performance is 50%. The arithmetic is exact for these invented counts; it is not an estimate of production performance without assumptions about sampling and future traffic. Describe who judged the replies and what “acceptable” meant.
Evidence must identify the tested system version, source snapshot, configuration, test cases, evaluation method, and unresolved failures. Check important slices of use, such as policy exceptions and supported languages, with enough relevant cases to interpret them. Sparse slices leave uncertainty; an empty slice is not a pass. Keep failures visible when reporting an average, and distinguish a suggested fix from a fix that has been retested.
The Model Cards for Model Reporting paper proposes documenting intended uses and performance under relevant conditions. Such a model card helps a reader understand a model’s evaluated scope. For the shop’s release decision, supplement component documentation with evidence about retrieval, permission enforcement, the interface, and staff review. A document describing a model does not certify the complete service.
Match sector requirements to the actual function
In lending, a useful fairness review investigates who is wrongly rejected or treated differently, rather than treating overall accuracy as sufficient. A difference in approval rates is a reason to investigate applicant circumstances, decision rules, and possible proxy variables; it does not by itself establish a legal violation. Where a statement of adverse-action reasons is required under U.S. Regulation B, it must identify specific principal reasons. A model-generated explanation or a hypothetical “change this input” suggestion does not automatically meet that requirement.
Financial controls also serve different purposes. Basel banking standards address matters such as capital, leverage, and liquidity; they are not a general AI approval standard. Know Your Customer (KYC) commonly refers to customer identification and verification within a wider due-diligence process that also assesses risk and maintains relevant monitoring. The FATF Recommendations provide international standards implemented through national measures. An AI identity check is only one possible component: failed matches need handling, sensitive documents need protection, and the institution must establish which requirements apply.
Medical AI can support imaging, clinical decisions, monitoring, or administrative work. A promising laboratory result does not demonstrate better patient outcomes in a particular clinic. Examine the intended patients, equipment, missed cases, false alarms, and how staff act on the output. Regulatory treatment depends on the function and jurisdiction: the FDA’s device-software policy distinguishes functions subject to its device oversight. The label “medical AI” alone does not determine approval requirements. Record the applicable basis and required evidence instead of copying a generic list of regulations.
Write a decision that can be operated
A decision record connects evidence to an outcome: decline, request more evidence, allow a bounded pilot, or approve the defined use with conditions. Record the deciding authority, rationale, version, allowed scope, open issues, owners, and review triggers. The organization must establish applicable legal and contractual requirements separately; an internal risk decision does not override them.
Illustrative decision for assistant version A
Status: no autonomous customer replies or refund actions
Unresolved: exception cases failed in the documented evaluation
Next work: improve exception routing; retest drafts and staff review
Possible pilot: staff-only drafting, only after pilot entry criteria pass
Before pilot: name owner, establish monitoring and a manual fallback
Re-review: changed policy sources, model, permissions, or intended use
The example deliberately does not approve a pilot merely because someone promises to monitor it. Define entry criteria and verify them first. During a pilot, the fallback must have capacity to handle the work when the assistant is unavailable. Human oversight also needs usable source evidence, time to inspect it, and a way to reject or escalate a draft. A required approval click alone provides little evidence of effective oversight.
Revisit approval when the evidence changes
After launch, track incorrect advice, staff corrections, escalations, complaints, and the intended benefit. An incident is an event requiring response, such as a draft containing another customer’s information. Provide a reporting channel, a responder, and predefined authority to restrict service. Preserve relevant evidence with appropriate access and retention controls, investigate affected cases, and decide what must be corrected before resuming. A failed release can require customer-facing correction as well as a software fix.
Review need not wait for an incident. A new model, changed refund rules, a new language, or the addition of payment tools can invalidate parts of earlier evidence. Link changes to the assumptions they affect and repeat the relevant evaluation and decision. On retirement, disable the workflow, remove permissions and dependencies where appropriate, and handle retained data under its lifecycle rules.
For a shared vocabulary, NIST AI RMF 1.0’s Core organizes work into Govern, Map, Measure, and Manage. These functions address governance, context, evaluation, and risk handling; they are not a one-time sequence. The framework overview describes voluntary use. Referring to it does not itself establish legal compliance or certify a system.
Test whether you can defend a decision
1. The shop reports 95% acceptable drafts and proposes automatic refunds. What evidence and decision are missing?
Solution
The score describes drafting on a particular test set, including only 50% success on the fictional exception cases. Refund execution changes the use and consequences. Define that new scope, evaluate action authorization and failure handling as well as decision quality, and obtain a decision from the authorized owner under the applicable requirements. The drafting result does not authorize the new use.
2. Every draft requires a staff approval click, but reviewers have no source documents and are rewarded only for speed. Has effective oversight been demonstrated?
Solution
No. The workflow records an action, not whether staff can detect and correct errors. Provide the evidence, time, competence, and authority needed for review, then observe whether the resulting process catches relevant failures. Include correction time and escalation workload when judging the benefit.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
