A practical field guide to AI governance internal audits — gold-standard frameworks, step-by-step methodology across 3 phases, 11 risk domains, common enterprise findings, and every resource you actually need. Based on IIA, NIST, ISO 42001, EU AI Act, ISACA, and Big 4 guidance through 2025.
gold-standard frameworks — 6 to know

Dog #1 says: "NIST AI RMF is the skeleton key. It doesn't tell you what to do — it tells you what questions to ask. Master the four functions and you can audit any AI system in any sector."
Published January 2023 by NIST (AI 100-1). Voluntary, technology-neutral, sector-agnostic. Not certifiable — organizations self-attest. Updated in July 2024 with AI 600-1, a Generative AI Profile adding 200+ actions across 12 GenAI-specific risk categories. The closest thing to a universal AI audit scaffold that exists.
GOVERN → Cross-cutting: AI policy, risk tolerance, accountability, third-party risk, incident response
MAP → Context: system boundaries, stakeholder impact, risk categorisation, AI inventory
MEASURE → Assessment: bias tests, red-teaming, performance drift, third-party evaluations
MANAGE → Treatment: risk response plans, escalation paths, decommissioning, post-incident reviews

Dog #2 says: "ISO 42001 is the ISO 27001 of AI. If you want a certificate to show your board, your customers, or your regulator, this is the one. The 38 Annex A controls are your audit checklist."
Published December 2023. World's first certifiable AI management system standard. Organizations undergo a two-stage external audit by an accredited Conformity Assessment Body (CAB) and receive a 3-year certificate with annual surveillance audits. KPMG was first Big 4 firm to achieve certification.
Clause 4: Context (AI use cases, interested parties, scope)
Clause 5: Leadership (AI policy — Clause 5.2 specific requirement)
Clause 6: Planning (risk assessment 6.1.2, impact assessment 6.1.4)
Clause 8: Operations (AI lifecycle, Annex A control implementation)
Clause 9: Performance evaluation (internal audit, management review)
Annex A: 38 controls across 9 objectives — A.5 Lifecycle, A.6 Data,
A.7 Transparency, A.8 Responsible Use, A.9 Third-Party

Dog #3 says: "The EU AI Act is GDPR for AI, but with sharper teeth. If your system makes decisions in employment, credit, education, or healthcare — and it affects EU persons — you're in scope. August 2026 is not optional."
Formally adopted May 2024, entered into force August 2024. Built on the OECD AI Principles and explicitly adopts the OECD definition of an AI system. Four risk tiers: Prohibited, High-Risk, Limited, Minimal. Penalties for high-risk violations: up to €15M or 3% of global turnover. For prohibited practices: up to €35M or 7%.
Art. 9: Risk management system — continuous, documented lifecycle process
Art. 10: Data governance — representative, bias-free, documented datasets
Art. 11: Technical documentation — full specs before market placement
Art. 12: Logging — automatic event logs retained for regulatory review
Art. 13: Transparency — instructions, capabilities, limitations to deployers
Art. 14: Human oversight — monitor, intervene, override, halt mechanisms
Art. 16: Provider obligations — QMS, EU database registration, CE marking

Dog #4 says: "The IIA AI Framework is what you take into the audit room. It won't make you a data scientist. But it will tell you what questions a non-data-scientist can credibly ask — and that's most of the job."
The IIA's authoritative AI audit guidance. Updated September 2024 to incorporate NIST AI RMF alignment, LLM risks, and generative AI considerations. Aligned to the 2024 Global Internal Audit Standards (effective Jan 2025). Organized around the IIA Three Lines Model — the closest thing to a GTAG for AI that exists.
LINE 1 — Governance (Board / Audit Committee):
Set AI risk appetite, ethical boundaries, oversight expectations
LINE 2 — Management (1st & 2nd line):
Build/use AI responsibly; own risk management, model inventory,
responsible AI procedures, incident response
LINE 3 — Internal Audit (3rd line):
Independent assurance & advisory over AI governance, risk, controls

Dog #5 says: "ISACA's toolkit is the most granular thing on the market. 127 control activities across 40 control objectives. If the IIA framework is a map, the ISACA toolkit is turn-by-turn navigation."
ISACA released a major 2025 white paper mapping COBIT to AI governance. The AI Audit and Risk Management Toolkit (40 control objectives, 127 control activities across bias, privacy, human-AI interaction, secure AI development) is available at $49/members, $99/non-members. ISACA's Journal published a proposed high-level AI audit approach in Volume 2, 2024.
5 AI Risk Categories:
Ethical Usage | Policy & Governance | Technology & Infrastructure
Operational & Organizational | Emerging & Strategic Risk
7 Governance Enablers:
Processes | Organizational Structures | Policies & Procedures
People / Skills / Competencies | Culture / Ethics / Behavior
Information Flows | Services / Infrastructure / Applications

Dog #6 says: "The Big 4 frameworks are where theory meets billable hours. Useful as benchmark documents — they're written for clients, so they're plain-English. Download their public guides before starting your audit planning."
All four major audit firms have published substantive AI governance guidance aligned to NIST AI RMF and ISO 42001. PwC launched the first independent AI assurance service (June 2025, AICPA standards). KPMG achieved ISO 42001 certification — the first Big 4 firm to do so.
step-by-step audit methodology — 3 phases
Obtain and review the organizational AI inventory from management. Identify shadow AI using IT asset scans, SaaS subscription audits, expense report reviews, and business unit interviews. Apply risk-tiering criteria: complexity, autonomy, decision impact, volume of affected individuals, regulatory exposure. Select audit universe — prioritize high-risk, high-impact systems.
Determine whether the audit team has sufficient AI technical literacy (data science, ML fundamentals, statistics). Engage subject matter experts (data scientists, ML engineers, ethicists) if needed. Identify applicable external standards for the audit scope: EU AI Act tier, NIST AI RMF profile, ISO 42001 clauses.
Define audit objectives across all 11 risk domains. Map objectives to the IIA AI Framework checklist and applicable regulatory requirements. Develop procedures for each objective: document review, interviews, technical testing, data analytics. Establish criteria for control design adequacy and operating effectiveness.
Determine which framework(s) apply (NIST AI RMF, ISO 42001, EU AI Act, sector-specific regulation such as SR 11-7 for financial services). Define what "adequate" looks like for each control — before you start fieldwork, not after. Document the criteria formally so findings can be rated against them.
Obtain AI policy, standards, and ethics principles documentation. Interview AI governance committee members or AI program office. Verify AI risk appetite is documented and approved at board/executive level. Assess AI strategy alignment to business objectives. No governance body = immediate finding.
Request the complete AI system inventory. Test completeness: cross-reference with IT asset management, vendor contracts, SaaS subscriptions, and business unit interviews. Verify each system has a documented owner, risk classification, and last review date. Identify any unregistered or shadow AI instances.
Review data lineage documentation for training datasets. Test data quality controls (completeness, accuracy, timeliness). Assess bias testing at data collection and labeling stages. Verify data privacy impact assessments were conducted before using personal data for training. EU AI Act Article 10 compliance check.
Review model development documentation: requirements, design, testing, validation, approval-to-deploy. Assess independence of model validation — was it done by the same team that built the model? Test that performance metrics are defined, measured, and within acceptable thresholds. Review fairness testing results across demographic groups. Model cards present and complete?
For high-stakes decisions (credit, hiring, healthcare, benefits): verify explanations can be generated for individual decisions. Test that human review mechanisms exist and are functioning — not rubber-stamping. Interview business users on their understanding of AI model limitations. Assess regulatory disclosure compliance (GDPR Art. 22, ECOA, EU AI Act Art. 13).
Verify human-in-the-loop procedures are documented and tested. Confirm operators can monitor, override, and halt the system (EU AI Act Art. 14 requirement). Test that operators are trained on the system's limitations. Check for automation bias — are humans actually interrogating outputs, or rubber-stamping?
Review AI system within the broader cybersecurity framework (penetration testing scope, vulnerability scanning). Test access controls on model weights, training pipelines, and inference APIs. Assess prompt injection protections for GenAI systems. Review model poisoning and adversarial input controls. Confirm AI-specific attack scenarios are included in IR playbooks.
Obtain vendor AI risk assessment documentation. Review AI-specific contract clauses: data use, model change notification, audit rights, sub-processor management. Test that vendor AI models are subject to equivalent governance as internal models. Map the full sub-processor chain — hidden AI services are a common gap.
Test model performance dashboards and alerting for model drift and data drift. Review AI incident log (past 12 months) for patterns and remediation quality. Assess incident response plan completeness. Review model retraining triggers and approval processes. Is there a decommissioning process for retired models?
Run adversarial prompting tests on GenAI systems. Perform boundary/stress tests on model performance edge cases. Conduct data-lineage tracing to verify documented provenance. Use data analytics for full-population testing of AI outputs for statistical bias signals. Deloitte recommends this as table-stakes for mature IA functions.
Verify a GenAI acceptable use policy exists and covers: data classification rules (which data can enter which tools), prohibited uses (client data, regulated data, credentials), output verification requirements, and consequences for policy violations. Sample GenAI tool usage logs or conduct anonymous employee surveys to assess actual compliance — not assumed compliance. Test whether the policy is actively enforced. Evidence of non-compliance should be escalated as a control failure, not simply an awareness gap. Acceptable use gaps are now a regulatory exposure under EU AI Act GPAI obligations (live August 2025).
For any AI system that takes multi-step autonomous actions (code execution, email sending, data retrieval, API calls) without per-step human approval: verify goal constraints and guardrails are formally defined and documented. Confirm all agent actions are logged with sufficient detail to reconstruct full decision chains post-incident. Test that human approval checkpoints exist for high-consequence actions. Assess the "blast radius" of each agentic system — what is the maximum damage a misaligned or compromised agent could cause? Verify leadership is aware of this exposure. Reference NIST Agentic AI Guidance (2025) and OWASP LLM Top 10 v2.0 — standard checklists do not yet adequately cover this domain.
Rate findings against pre-defined criteria (control design adequacy vs. operating effectiveness). Use risk-based severity: Critical / High / Medium / Low. Map each finding to a risk domain. The absence of governance structures is itself a Critical finding — do not soften it. Present maturity ratings across domains, not just a list of issues.
Present findings to AI system owners, AI governance committee, CAE, and Audit Committee. Include a heat map of risk domain coverage and maturity ratings. Recommend remediation with target dates and responsible parties. Frame regulatory exposure concretely — "this gap is a potential EU AI Act non-conformance ahead of the August 2026 deadline" lands differently than abstract risk language.
Track remediation of high/critical findings within 90 days. Validate corrective actions through evidence review or re-testing — do not close based on management representation alone. For EU AI Act gaps: track progress against the August 2026 deadline explicitly in the follow-up register.
11 AI audit risk domains
The foundation. If there's no AI strategy, no governance body, and no ethics principles, everything downstream is built on air. This domain is often the first finding — and often the most significant.
You can't govern what you can't see. An incomplete AI inventory is the most common and consequential gap — shadow AI (unauthorized tools used by employees) and AI embedded in SaaS platforms create invisible risk. 56–80% of employees use unauthorized AI tools.
Bad data makes bad models. The EU AI Act Article 10 and ISO 42001 Annex A.6 both require documented, representative, bias-free training datasets. Data governance for AI extends beyond the model — it includes consent for training use, retention of training data vs. model weights, and cross-border transfer controls.
SR 11-7 (Federal Reserve / OCC model risk guidance) applies to AI/ML models in financial services — but the principles are sound for any sector. Independent validation, documented conceptual soundness, and ongoing performance monitoring are the three pillars. Black-box models deployed without independent validation are a critical finding.
Explainability is transitioning from best practice to legal requirement. ECOA/FCRA (US), GDPR Art. 22 (EU), and EU AI Act Art. 13 all require meaningful explanations for automated decisions affecting individuals. Adverse action notices generated by black-box models are a major regulatory exposure.
EU AI Act Article 14 requires that high-risk AI systems be designed so humans can monitor, intervene, override, and halt them. "Human oversight" that exists on paper but consists of rubber-stamping AI outputs is not compliance. Test whether human reviewers actually exercise judgment.
AI systems face attack vectors that traditional cybersecurity frameworks don't fully cover. AI-related security incidents surged 56.4% in 2024–2025. Prompt injection attacks exploited legitimate GenAI tools at 90+ organizations to steal credentials. OWASP Top 10 for LLMs (2025 edition) is the practitioner reference.
Most organizations use more AI from vendors than they build themselves — but vendor AI often receives far less scrutiny than internal models. Sub-processor chains (vendor uses another AI provider) create hidden risk. AI features embedded in approved SaaS tools (Slack, Salesforce, Microsoft 365) frequently bypass the AI intake process entirely.
The regulatory landscape is accelerating. EU AI Act high-risk obligations are live for enforcement August 2, 2026. Prohibited practices have been banned since February 2025. Sector-specific rules (SR 11-7 for financial services, FDA AI guidance for medical devices, EEOC for employment) layer on top of cross-cutting frameworks.
AI systems degrade over time as real-world data drifts away from training data. Model drift monitoring is not optional — it's how you detect performance degradation before it becomes a customer harm event or a regulatory incident. The absence of drift monitoring in production is a High finding.
Agentic AI (autonomous multi-step AI systems that take actions in the world) is the fastest-evolving audit frontier in 2025. These systems can make sequences of decisions and take real-world actions without per-step human review. The audit challenges are traceability, goal alignment, and guardrail governance. Standard audit techniques don't yet cover this domain adequately.
top findings from enterprise AI audits
EU AI Act compliance timeline
key resources — updated june 2025
what's new — key guidance since may 2025
questions to ask your organisation — internal AI governance audit

Dog #1 says: "If governance questions get blank stares, you already have your first finding. The absence of an answer is the answer."

Dog #2 says: "Ask for the inventory first. Then go find what's not in it. The gap between those two lists is your audit."

Dog #3 says: "The question nobody thinks to ask: was this personal data collected with consent for AI training specifically — not just 'we had it and needed it'?"

Dog #4 says: "The question that exposes everything: who validated this model before it went live — and were they on the same team that built it?"

Dog #5 says: "August 2, 2026 is not a vague future date — it is 14 months away. No conformity assessment plan means you have a regulatory exposure finding right now, today."

Dog #6 says: "Are your human reviewers actually reviewing — or just approving? One of these is compliance. The other is theatre."

Dog #1 says: "'We bought it from a vendor' is not a control. It's a statement that the work hasn't been done yet."

Dog #2 says: "AI-enabled fraud is not a future scenario — voice cloning BEC and AI-generated phishing are live threats right now. Has your SOC run a tabletop on this yet?"

Dog #3 says: "The most dangerous question in AI governance right now: what is the maximum autonomous action your AI can take without human approval — and does leadership know the answer?"

Dog #4 says: "You're auditing AI while AI is being used to audit. The function that ignores its own AI governance gap is the least credible auditor in the room."
golden rules for auditing artificially intelligent nerds