DECISION TOOLKIT · GOVERNANCE OPERATING MODEL

Scale control intensity with consequence

Risk is not a property of a model name. It is the consequence of a workflow, the authority granted to the system and the organization’s ability to detect and reverse an error.

01Risk tier → minimum control posture
01Assist

drafting, search, low-impact internal work

approved tool · user accountability · light logging
02Inform

professional judgment or external content

evaluation set · reviewer · traceable output
03Act

writes to systems or affects customers

least privilege · validated calls · confirmation
04Critical

rights, safety, finance or regulated decisions

formal approval · independent assurance · incident drill

Fast track low-risk uses; reserve the strongest governance for the workflows that can create material harm.

02Evidence is the control plane
01Inventory
02Classify
03Control
04Assure
05Respond

AI governance fails when it remains a policy document. Teams need to know which systems exist, what they may do, which controls apply and what evidence is required before a higher-risk use enters production.

An operating framework should make the safe path clear. It should also remain proportionate: if every use receives the same heavy process, governance will drive work into the shadows.

Establish the system inventory

Begin with AI systems, not only approved vendors. Record the business owner, users, workflow, data categories, models, integrations, deployment boundary and decision impact. Include embedded AI in existing software and unsanctioned tools where material.

This inventory becomes the control plane for risk classification, assurance and incident response. It should evolve through procurement and engineering workflows rather than depend on a yearly survey.

Classify risk at the workflow level

The same model can support low- and high-risk activities. Classify the actual use based on data sensitivity, user exposure, autonomy, reversibility, decision consequence and the ability to detect error.

One practical structure uses four tiers:

  1. Assist — low-impact internal productivity with human responsibility retained.
  2. Inform — outputs influence professional judgment or external content.
  3. Act — the system takes actions in connected systems or affects customers.
  4. Critical — material rights, safety, finance or regulated decisions may be affected.

The tier determines approval, evaluation, monitoring and escalation requirements.

Design controls across the lifecycle

Controls should cover more than model behavior. They include purpose limitation, lawful data access, identity and permissions, supplier terms, prompt and context handling, evaluation, human oversight, logging, incident response and change management.

Control domainMinimum evidence
Datacategories, source, access basis, retention and residency
Securitythreat model, trust boundaries, permissions and abuse tests
Qualityrepresentative evaluations and acceptance thresholds
Human oversightnamed decision owner and intervention path
Operationsmonitoring, rollback, incident and change process
Supplierservice terms, sub-processors, model changes and exit plan

Treat generative AI as a new attack surface

Prompt injection, tool misuse, data exfiltration, insecure retrieval, over-permissioned agents and poisoned context are system risks. They cannot be solved by telling users to be careful.

Threat-model the full chain: user, interface, orchestration, tools, retrieval, model, output consumers and external services. Reduce privileges, validate tool calls, isolate untrusted content and create explicit confirmation points for consequential actions.

Make evidence reusable

Assurance becomes expensive when every project invents its own documents. Define shared templates for system cards, data-flow maps, evaluation reports, risk acceptance and production readiness. Platform controls should produce evidence automatically where possible.

Governance forums then focus on exceptions and decisions, not collecting screenshots.

Assign decision rights

Every system needs a business owner accountable for its purpose and outcome, a technical owner accountable for implementation and operations, and a risk owner accountable for control acceptance. Legal, security, data protection and procurement contribute expertise; they should not become the default owner of every AI outcome.

The final test is operational: can a team tell what it is allowed to do, can a leader see which risks were accepted, and can the organization respond quickly when behavior or suppliers change? If yes, governance has become infrastructure for execution rather than an obstacle beside it.

Use a risk model leaders can actually operate

The four tiers are deliberately simple. Their purpose is not to produce a perfect legal taxonomy; it is to create a common language between the business owner, engineering, security, legal and the executive who accepts residual risk. A classification that requires a specialist to explain it will not survive a procurement queue or an incident at 18:00.

Use six questions to place a workflow. How sensitive is the data? Who is exposed to the output? What authority does the system have? How reversible is an error? How detectable is a bad result? What is the consequence if the workflow is wrong at scale? Score each question from one to four, then let the highest consequence dimension set the tier. This “highest-water mark” rule is safer than averaging away a material exposure.

Risk questionLow signalHigh signalWhat it changes
Data sensitivitypublic or internal reference materialpersonal, confidential, privileged or regulated datadata boundary, retention, transfer and access review
System authoritydrafts and recommendswrites, sends, approves or triggers downstream workleast privilege, confirmation and rollback
Decision consequenceeasy to correct, no external impactrights, safety, money, eligibility or reputationhuman oversight, assurance and escalation
Error detectabilityvisible immediately to a competent userplausible error may pass unnoticedevaluation depth, sampling and monitoring
Exposureone trained usercustomers, employees or third parties at scaletransparency, incident response and audit evidence

The result should be a workflow record, not a model label. “Claude is low risk” or “the internal GPT is approved” are not meaningful statements. “The claims assistant drafts responses from approved policy sources, never sends externally, and requires a named reviewer” is a governable statement.

Build a control library, not a control burden

Controls should be reusable objects that teams can attach to a workflow. The library can be small at first, provided each control has an owner, an implementation pattern, an evidence output and a test frequency.

Identity and access. Provision through the corporate identity provider, separate administrators from users, map groups to data classes and remove access when employment or role changes. For connected agents, the service identity should be narrower than the human identity wherever possible. “The user can access it” is not sufficient justification for “the agent may act on it.”

Data boundary. Declare which sources are in scope, what is excluded, the retention period, the processing location and the permitted output destinations. Retrieval should enforce source permissions at query time rather than rely on a prompt that says “do not reveal confidential content.”

Quality and evaluation. Maintain a versioned set of representative tasks: ordinary cases, ambiguous cases, adversarial cases, multilingual cases and known historical failures. Store not only a pass rate but the reason for failure, severity, detectability and whether a human caught it.

Action control. Separate read, propose and execute permissions. Validate tool arguments against a schema, apply business rules outside the model and require confirmation for irreversible actions. The model can propose a payment, a customer communication or a deletion; the system should decide whether the action is allowed.

Observability. Capture enough trace to reconstruct what happened without creating disproportionate employee surveillance. A useful trace includes workflow identifier, actor, model or version, tools invoked, policy outcome, reviewer decision, cost and latency. Retention should follow the data class and the incident need.

Make governance a service with explicit service levels

Governance teams create shadow AI when they behave like a committee that only says no. A better operating model offers a published intake path, reusable patterns and response times. Low-risk Assist uses can follow a fast track: approved environment, standard training, minimum logging and self-service registration. Inform and Act uses require evidence packs and named reviewers. Critical uses require a formal decision forum and an independent challenge.

The service should publish the questions it will ask before a team starts building. That single intervention changes the quality of projects: owners collect data-flow information earlier, engineers avoid impossible patterns and procurement understands why an enterprise contract is required.

Run a monthly operating review with four fixed sections:

  1. Inventory movement — new systems, changed systems, retired systems and material unapproved uses.
  2. Control health — access exceptions, evaluation drift, overdue reviews, logging coverage and vendor changes.
  3. Incidents and near misses — what happened, what was detected, which control failed and what will change.
  4. Decisions required — risk acceptance, escalation, restriction, investment or retirement.

Quarterly, review whether the tiering model is still proportionate. A system may move from Assist to Inform because its users begin relying on it for external content. An Act workflow may become Critical because the business connects it to a new decision process. Classification is a living decision, not a launch milestone.

The executive test: can the organization show its work?

The board-level question is not whether a company has an AI policy. It is whether management can show a controlled path from use to consequence. The minimum executive pack should contain the inventory by risk tier, the top open exceptions, the material vendor and model changes, the incidents and near misses, the cost of controls and the decisions requested.

Three failure modes should trigger immediate attention. First, the inventory is smaller than the number of tools employees can access. Second, the policy forbids a behavior but the approved platform makes it the easiest path. Third, the organization cannot reproduce why a higher-risk use was approved. Each is a design failure, not a training failure.

The practical standard is demanding but clear: every material use has an owner; every owner knows the permitted authority; every higher-risk workflow has a repeatable evidence pack; every residual risk is accepted by a named person; and every incident creates a change in the system, the control or the decision. That is how governance becomes a production capability.