rules for artificially intelligent nerds

A practical field guide to AI governance internal audits — gold-standard frameworks, step-by-step methodology across 3 phases, 11 risk domains, common enterprise findings, and every resource you actually need. Based on IIA, NIST, ISO 42001, EU AI Act, ISACA, and Big 4 guidance through 2025.

83%
of organizations have no enterprise-wide AI policy (IIA / Wolters Kluwer 2024)
Aug 2026
EU AI Act high-risk system obligations become enforceable — the clock is ticking
60%+
of Chief Audit Executives cite insufficient AI skills within their audit teams
56–80%
of employees use unauthorized AI tools — shadow AI is the #1 inventory gap

gold-standard frameworks — 6 to know

01 dog
NIST AI Risk Management Framework (AI RMF 1.0)
voluntary, U.S.-origin, globally adopted — the practitioner's map
dog

Dog #1 says: "NIST AI RMF is the skeleton key. It doesn't tell you what to do — it tells you what questions to ask. Master the four functions and you can audit any AI system in any sector."

Published January 2023 by NIST (AI 100-1). Voluntary, technology-neutral, sector-agnostic. Not certifiable — organizations self-attest. Updated in July 2024 with AI 600-1, a Generative AI Profile adding 200+ actions across 12 GenAI-specific risk categories. The closest thing to a universal AI audit scaffold that exists.

four core functions
GOVERN → Cross-cutting: AI policy, risk tolerance, accountability, third-party risk, incident response MAP → Context: system boundaries, stakeholder impact, risk categorisation, AI inventory MEASURE → Assessment: bias tests, red-teaming, performance drift, third-party evaluations MANAGE → Treatment: risk response plans, escalation paths, decommissioning, post-incident reviews

key audit uses

  • GOVERN GV-1: Does an AI risk policy exist?
  • GOVERN GV-2.3: Has leadership declared risk tolerances?
  • GOVERN GV-5: Third-party / vendor AI risk management
  • MAP: Is every AI system catalogued with purpose + limitations?
  • MEASURE: Evidence of bias testing, red-teaming, drift monitoring
  • MANAGE: Documented risk treatment decisions and escalation paths

2024–2025 updates

  • AI 600-1 (Jul 2024): GenAI Profile — 12 risk categories
  • 12 categories incl. hallucination, prompt injection, IP, privacy
  • Agentic AI Profile in development (CSA Lab Space 2025)
  • 58-control audit checklist available (AIGL)
  • Maps well to ISO 42001 PDCA — use together

official resources

nist.gov/itl/ai-risk-management-framework airc.nist.gov/airmf-resources/playbook NIST AI 600-1 GenAI Profile PDF
02 dog
ISO/IEC 42001:2023 — AI Management System Standard
certifiable, international, PDCA-based — the gold standard for proof
dog

Dog #2 says: "ISO 42001 is the ISO 27001 of AI. If you want a certificate to show your board, your customers, or your regulator, this is the one. The 38 Annex A controls are your audit checklist."

Published December 2023. World's first certifiable AI management system standard. Organizations undergo a two-stage external audit by an accredited Conformity Assessment Body (CAB) and receive a 3-year certificate with annual surveillance audits. KPMG was first Big 4 firm to achieve certification.

10 clauses + 38 annex A controls
Clause 4: Context (AI use cases, interested parties, scope) Clause 5: Leadership (AI policy — Clause 5.2 specific requirement) Clause 6: Planning (risk assessment 6.1.2, impact assessment 6.1.4) Clause 8: Operations (AI lifecycle, Annex A control implementation) Clause 9: Performance evaluation (internal audit, management review) Annex A: 38 controls across 9 objectives — A.5 Lifecycle, A.6 Data, A.7 Transparency, A.8 Responsible Use, A.9 Third-Party

Stage 1 audit (documentation)

  • Scope definition and AI policy (5.2)
  • Risk assessment methodology
  • AI impact assessments (6.1.4)
  • Statement of Applicability (SoA)
  • Internal audit records
  • Management review documentation

most common audit gaps

  • Documentation written after deployment (not before)
  • No fairness testing records across demographics
  • Weak third-party AI vendor controls
  • Human oversight checkpoints not evidenced
  • Impact assessments not updated when models change

official resources

iso.org/standard/42001 ISO 42001 Annex A 38 Controls Guide (ISAuditr) Cloud Security Alliance lessons learned (May 2025)
03 dog
EU AI Act (2024) — world's first binding AI regulation
risk-tiered, enforceable, fines up to 7% of global revenue
dog

Dog #3 says: "The EU AI Act is GDPR for AI, but with sharper teeth. If your system makes decisions in employment, credit, education, or healthcare — and it affects EU persons — you're in scope. August 2026 is not optional."

Formally adopted May 2024, entered into force August 2024. Built on the OECD AI Principles and explicitly adopts the OECD definition of an AI system. Four risk tiers: Prohibited, High-Risk, Limited, Minimal. Penalties for high-risk violations: up to €15M or 3% of global turnover. For prohibited practices: up to €35M or 7%.

high-risk AI — articles 9–17 checklist
Art. 9: Risk management system — continuous, documented lifecycle process Art. 10: Data governance — representative, bias-free, documented datasets Art. 11: Technical documentation — full specs before market placement Art. 12: Logging — automatic event logs retained for regulatory review Art. 13: Transparency — instructions, capabilities, limitations to deployers Art. 14: Human oversight — monitor, intervene, override, halt mechanisms Art. 16: Provider obligations — QMS, EU database registration, CE marking

high-risk AI categories (Annex III)

  • Biometric identification & categorisation
  • Critical infrastructure (energy, water, transport)
  • Education & vocational training
  • Employment, recruitment, promotion decisions
  • Credit scoring, insurance, essential services access
  • Law enforcement & border control
  • Administration of justice

internal auditor checklist

  • AI system inventory with risk tier classification
  • Prohibited practice check (Feb 2025 obligation live)
  • Conformity assessment track determined (self vs. notified body)
  • Training data bias documentation in place
  • Human-in-the-loop mechanisms designed and tested
  • Post-market monitoring plan documented
  • EU AI database registration completed

key dates

Feb 2025: Prohibited practices banned Aug 2025: GPAI model obligations Aug 2026: High-risk system enforcement artificialintelligenceact.eu
04 dog
IIA AI Auditing Framework (Sept 2024 Update)
the internal auditor's primary reference — three lines model, practitioner checklist
dog

Dog #4 says: "The IIA AI Framework is what you take into the audit room. It won't make you a data scientist. But it will tell you what questions a non-data-scientist can credibly ask — and that's most of the job."

The IIA's authoritative AI audit guidance. Updated September 2024 to incorporate NIST AI RMF alignment, LLM risks, and generative AI considerations. Aligned to the 2024 Global Internal Audit Standards (effective Jan 2025). Organized around the IIA Three Lines Model — the closest thing to a GTAG for AI that exists.

three lines model for AI
LINE 1 — Governance (Board / Audit Committee): Set AI risk appetite, ethical boundaries, oversight expectations LINE 2 — Management (1st & 2nd line): Build/use AI responsibly; own risk management, model inventory, responsible AI procedures, incident response LINE 3 — Internal Audit (3rd line): Independent assurance & advisory over AI governance, risk, controls

framework structure (4 parts)

  • Part 1: AI overview (history, types, GenAI landscape)
  • Part 2: Getting started (scope, skills, planning)
  • Part 3: Detailed AI auditing framework by three lines
  • Part 4: Practitioner's guide, checklist, glossary

IIA 2025 data points

  • Digital disruption: top-5 risk for 48% of CAEs (up 9pts YoY)
  • GenAI use in IA: doubled 15% → 40% in one year
  • 4 in 10 IA functions not prepared for AI-enabled fraud
  • Only 4% of CAEs report substantial AI progress in IA itself
  • 60%+ of CAEs cite AI skills gap

official resources

IIA AI Auditing Framework PDF (Sept 2024) theiia.org AI Knowledge Center IIA "Auditing AI" Hands-On Course
05 dog
ISACA / COBIT for AI Governance (2025)
40 control objectives, 127 control activities, 5 risk categories, 7 governance enablers
dog

Dog #5 says: "ISACA's toolkit is the most granular thing on the market. 127 control activities across 40 control objectives. If the IIA framework is a map, the ISACA toolkit is turn-by-turn navigation."

ISACA released a major 2025 white paper mapping COBIT to AI governance. The AI Audit and Risk Management Toolkit (40 control objectives, 127 control activities across bias, privacy, human-AI interaction, secure AI development) is available at $49/members, $99/non-members. ISACA's Journal published a proposed high-level AI audit approach in Volume 2, 2024.

COBIT for AI — 5 risk categories × 7 enablers
5 AI Risk Categories: Ethical Usage | Policy & Governance | Technology & Infrastructure Operational & Organizational | Emerging & Strategic Risk 7 Governance Enablers: Processes | Organizational Structures | Policies & Procedures People / Skills / Competencies | Culture / Ethics / Behavior Information Flows | Services / Infrastructure / Applications

ISACA AI toolkit coverage

  • Algorithm audits: key control considerations (2024)
  • Shadow AI auditing guide (2025)
  • Third-party AI risk: 6-step framework (2025)
  • Adversarial ML threat guidance (2025)
  • Agentic AI auditing challenge (2025)
  • XAI assurance: new era guidance (2025)

key ISACA resources

  • COBIT for AI Governance white paper (2025)
  • AI Algorithm Audits: Key Controls (ISACA Journal 2024)
  • DTEF (Digital Trust Ecosystem Framework) for bias/XAI
  • ISACA Now Blog: 5 Ways IT Auditors Can Use AI
  • AI Audit and Risk Toolkit ($49 / $99)

official resources

isaca.org AI Audit Toolkit ISACA COBIT AI Governance white paper (2025) ISACA Journal Volume 2 2024
06 dog
Big 4 AI Audit Frameworks
Deloitte Trustworthy AI · PwC Assurance for AI · EY IA Adaptation · KPMG Trusted AI
dog

Dog #6 says: "The Big 4 frameworks are where theory meets billable hours. Useful as benchmark documents — they're written for clients, so they're plain-English. Download their public guides before starting your audit planning."

All four major audit firms have published substantive AI governance guidance aligned to NIST AI RMF and ISO 42001. PwC launched the first independent AI assurance service (June 2025, AICPA standards). KPMG achieved ISO 42001 certification — the first Big 4 firm to do so.

Deloitte Trustworthy AI

  • 5 pillars: transparency, fairness, privacy, reliability, accountability
  • Audit techniques: adversarial prompting, boundary/stress tests, data-lineage tracing
  • Hot Topics 2025: inventory + agentic AI = top governance gaps
  • Free public guide: "Internal Audit's Role in Strengthening AI Governance"

KPMG Trusted AI Framework

  • 4 integrated enablers: Control Framework, Responsible AI Toolkit, Compliance Heatmap, AI Governance Blueprint
  • Heatmap covers: EU AI Act + ISO 42001 + NIST AI RMF cross-mapping
  • Published "Illustrative AI Risk and Controls Guide" (free download)
  • ISO 42001 certified — world's first Big 4 certification

PwC Assurance for AI (Jun 2025)

  • First independent AI assurance service under AICPA standards
  • Aligned to NIST AI RMF + ISO 42001
  • Responsible AI Toolkit: readiness assessments + controls design
  • 2025 survey: "Responsible AI — From Policy to Practice"

EY Internal Audit Adaptation

  • "How Internal Audit Can Adapt to AI" — public guidance
  • Covers: AI strategy maturity, model risk, bias, data integrity
  • Frames IAF role in ongoing AI governance

step-by-step audit methodology — 3 phases

P1 dog
Phase 1 — Pre-Engagement Planning
risk assessment · scoping · capability check · audit program
1.1

Risk Assessment & Scoping

Obtain and review the organizational AI inventory from management. Identify shadow AI using IT asset scans, SaaS subscription audits, expense report reviews, and business unit interviews. Apply risk-tiering criteria: complexity, autonomy, decision impact, volume of affected individuals, regulatory exposure. Select audit universe — prioritize high-risk, high-impact systems.

1.2

Capability Assessment

Determine whether the audit team has sufficient AI technical literacy (data science, ML fundamentals, statistics). Engage subject matter experts (data scientists, ML engineers, ethicists) if needed. Identify applicable external standards for the audit scope: EU AI Act tier, NIST AI RMF profile, ISO 42001 clauses.

1.3

Audit Objectives & Program Development

Define audit objectives across all 11 risk domains. Map objectives to the IIA AI Framework checklist and applicable regulatory requirements. Develop procedures for each objective: document review, interviews, technical testing, data analytics. Establish criteria for control design adequacy and operating effectiveness.

1.4

Establish the Audit Criteria Baseline

Determine which framework(s) apply (NIST AI RMF, ISO 42001, EU AI Act, sector-specific regulation such as SR 11-7 for financial services). Define what "adequate" looks like for each control — before you start fieldwork, not after. Document the criteria formally so findings can be rated against them.

P2 dog
Phase 2 — Fieldwork & Control Testing
10 testing workstreams across all AI risk domains
2.1

Governance Review

Obtain AI policy, standards, and ethics principles documentation. Interview AI governance committee members or AI program office. Verify AI risk appetite is documented and approved at board/executive level. Assess AI strategy alignment to business objectives. No governance body = immediate finding.

2.2

AI Inventory Validation

Request the complete AI system inventory. Test completeness: cross-reference with IT asset management, vendor contracts, SaaS subscriptions, and business unit interviews. Verify each system has a documented owner, risk classification, and last review date. Identify any unregistered or shadow AI instances.

2.3

Data Governance Testing

Review data lineage documentation for training datasets. Test data quality controls (completeness, accuracy, timeliness). Assess bias testing at data collection and labeling stages. Verify data privacy impact assessments were conducted before using personal data for training. EU AI Act Article 10 compliance check.

2.4

Model Development & Validation Lifecycle

Review model development documentation: requirements, design, testing, validation, approval-to-deploy. Assess independence of model validation — was it done by the same team that built the model? Test that performance metrics are defined, measured, and within acceptable thresholds. Review fairness testing results across demographic groups. Model cards present and complete?

2.5

Explainability & Transparency Testing

For high-stakes decisions (credit, hiring, healthcare, benefits): verify explanations can be generated for individual decisions. Test that human review mechanisms exist and are functioning — not rubber-stamping. Interview business users on their understanding of AI model limitations. Assess regulatory disclosure compliance (GDPR Art. 22, ECOA, EU AI Act Art. 13).

2.6

Human Oversight Mechanisms

Verify human-in-the-loop procedures are documented and tested. Confirm operators can monitor, override, and halt the system (EU AI Act Art. 14 requirement). Test that operators are trained on the system's limitations. Check for automation bias — are humans actually interrogating outputs, or rubber-stamping?

2.7

Cybersecurity & Adversarial Risk

Review AI system within the broader cybersecurity framework (penetration testing scope, vulnerability scanning). Test access controls on model weights, training pipelines, and inference APIs. Assess prompt injection protections for GenAI systems. Review model poisoning and adversarial input controls. Confirm AI-specific attack scenarios are included in IR playbooks.

2.8

Third-Party AI Vendor Risk

Obtain vendor AI risk assessment documentation. Review AI-specific contract clauses: data use, model change notification, audit rights, sub-processor management. Test that vendor AI models are subject to equivalent governance as internal models. Map the full sub-processor chain — hidden AI services are a common gap.

2.9

Monitoring & Incident Management

Test model performance dashboards and alerting for model drift and data drift. Review AI incident log (past 12 months) for patterns and remediation quality. Assess incident response plan completeness. Review model retraining triggers and approval processes. Is there a decommissioning process for retired models?

2.10

Technical Testing (where capability exists)

Run adversarial prompting tests on GenAI systems. Perform boundary/stress tests on model performance edge cases. Conduct data-lineage tracing to verify documented provenance. Use data analytics for full-population testing of AI outputs for statistical bias signals. Deloitte recommends this as table-stakes for mature IA functions.

2.11

GenAI Acceptable Use & Employee Policy Compliance

Verify a GenAI acceptable use policy exists and covers: data classification rules (which data can enter which tools), prohibited uses (client data, regulated data, credentials), output verification requirements, and consequences for policy violations. Sample GenAI tool usage logs or conduct anonymous employee surveys to assess actual compliance — not assumed compliance. Test whether the policy is actively enforced. Evidence of non-compliance should be escalated as a control failure, not simply an awareness gap. Acceptable use gaps are now a regulatory exposure under EU AI Act GPAI obligations (live August 2025).

2.12

Agentic AI & Autonomous System Governance

For any AI system that takes multi-step autonomous actions (code execution, email sending, data retrieval, API calls) without per-step human approval: verify goal constraints and guardrails are formally defined and documented. Confirm all agent actions are logged with sufficient detail to reconstruct full decision chains post-incident. Test that human approval checkpoints exist for high-consequence actions. Assess the "blast radius" of each agentic system — what is the maximum damage a misaligned or compromised agent could cause? Verify leadership is aware of this exposure. Reference NIST Agentic AI Guidance (2025) and OWASP LLM Top 10 v2.0 — standard checklists do not yet adequately cover this domain.

P3 dog
Phase 3 — Reporting & Follow-Up
rating findings · communicating results · tracking remediation
3.1

Rating & Criteria

Rate findings against pre-defined criteria (control design adequacy vs. operating effectiveness). Use risk-based severity: Critical / High / Medium / Low. Map each finding to a risk domain. The absence of governance structures is itself a Critical finding — do not soften it. Present maturity ratings across domains, not just a list of issues.

3.2

Communicating Results

Present findings to AI system owners, AI governance committee, CAE, and Audit Committee. Include a heat map of risk domain coverage and maturity ratings. Recommend remediation with target dates and responsible parties. Frame regulatory exposure concretely — "this gap is a potential EU AI Act non-conformance ahead of the August 2026 deadline" lands differently than abstract risk language.

3.3

Follow-Up & Validation

Track remediation of high/critical findings within 90 days. Validate corrective actions through evidence review or re-testing — do not close based on management representation alone. For EU AI Act gaps: track progress against the August 2026 deadline explicitly in the follow-up register.

11 AI audit risk domains

01 dog
AI Strategy & Governance
policy, risk appetite, ethics, accountability structures

The foundation. If there's no AI strategy, no governance body, and no ethics principles, everything downstream is built on air. This domain is often the first finding — and often the most significant.

audit questions

  • Is there a formal AI strategy aligned to business objectives?
  • Is there a board-level or executive AI governance body / CAIO?
  • Are AI ethical principles documented and operationalized (not just posted on a website)?
  • Is there an AI risk appetite statement?
  • Are AI policies formally approved and communicated to all staff?

red flags

  • No AI governance body with defined authority
  • Ethics principles exist only on paper — no evidence of application
  • AI risk appetite not formally set by leadership
  • Policy exists but has not been communicated or trained
  • No one owns AI governance enterprise-wide
02 dog
AI Inventory & Classification
shadow AI, risk tiering, ownership assignment

You can't govern what you can't see. An incomplete AI inventory is the most common and consequential gap — shadow AI (unauthorized tools used by employees) and AI embedded in SaaS platforms create invisible risk. 56–80% of employees use unauthorized AI tools.

audit questions

  • Does a comprehensive AI model inventory exist covering all deployed models?
  • Does it include SaaS-embedded AI, RPA bots with LLM components, and pilots?
  • Are systems classified by risk tier?
  • Is ownership and accountability assigned to each AI system?
  • When was the inventory last audited for completeness?

discovery techniques

  • Network traffic analysis and DLP log review
  • Browser history audits and expense report review
  • Employee surveys (anonymous)
  • SaaS subscription audit — scan for AI-enabled tools
  • Cross-reference IT asset list vs. AI inventory
  • Interview every business unit lead
03 dog
Data Governance & Quality
lineage, bias at source, consent, retention, cross-border transfers

Bad data makes bad models. The EU AI Act Article 10 and ISO 42001 Annex A.6 both require documented, representative, bias-free training datasets. Data governance for AI extends beyond the model — it includes consent for training use, retention of training data vs. model weights, and cross-border transfer controls.

audit questions

  • Is training data lineage documented to source?
  • Are data quality controls in place before data enters AI pipelines?
  • Was training data collected with appropriate consent for AI use?
  • Are data retention schedules specific to AI training data and model weights?
  • Are cross-border data transfers for AI training/inference assessed?

benchmark frameworks

  • DAMA DMBOK for data governance structure
  • EU AI Act Art. 10 — data representativeness requirement
  • ISO 8000 — data quality standard
  • ARMA GARP — retention policy governance
  • GDPR Art. 5 — purpose limitation for training data
04 dog
Model Development & Validation
lifecycle controls, independent validation, bias testing, model drift

SR 11-7 (Federal Reserve / OCC model risk guidance) applies to AI/ML models in financial services — but the principles are sound for any sector. Independent validation, documented conceptual soundness, and ongoing performance monitoring are the three pillars. Black-box models deployed without independent validation are a critical finding.

audit questions

  • Is there a documented AI/ML development lifecycle?
  • Are models validated independently before deployment?
  • Are performance metrics (accuracy, precision, recall, F1) defined and measured?
  • Are models tested for bias across demographic groups?
  • Are model cards / technical documentation complete?
  • Is there a model drift detection process?

fairness metrics to check

  • Demographic parity: equal outcomes across groups?
  • Equalized odds: equal false positive/negative rates?
  • Predictive parity: equal accuracy rates across groups?
  • Individual fairness: similar inputs → similar outputs?
  • Counterfactual fairness: would outcome change if group attribute changed?
05 dog
Explainability & Transparency
XAI techniques, adverse action notices, regulatory disclosure obligations

Explainability is transitioning from best practice to legal requirement. ECOA/FCRA (US), GDPR Art. 22 (EU), and EU AI Act Art. 13 all require meaningful explanations for automated decisions affecting individuals. Adverse action notices generated by black-box models are a major regulatory exposure.

audit questions

  • Can individual AI decisions be explained to affected parties?
  • Are XAI techniques (SHAP, LIME, attention maps) used appropriately?
  • Do adverse action notices meet regulatory specificity standards?
  • Are human reviewers actually interrogating model outputs — or rubber-stamping?
  • Can internal audit and regulators access model logic and decision traces?

regulatory requirements

  • ECOA / FCRA: specific reasons for adverse credit action
  • GDPR Art. 22: right to explanation for automated decisions
  • EU AI Act Art. 13: instructions, capabilities, limitations to deployers
  • NYC Local Law 144: annual bias audit requirement (hiring)
  • Colorado SB 205: algorithmic fairness (effective Feb 2026)
06 dog
Human Oversight & Controls
human-in-the-loop, automation bias, override mechanisms

EU AI Act Article 14 requires that high-risk AI systems be designed so humans can monitor, intervene, override, and halt them. "Human oversight" that exists on paper but consists of rubber-stamping AI outputs is not compliance. Test whether human reviewers actually exercise judgment.

audit questions

  • Are human-in-the-loop controls designed and operating for high-risk decisions?
  • Is there protection against automation bias?
  • Are override mechanisms documented, tested, and actually used?
  • Are AI-generated outputs reviewed before consequential actions?
  • Are operators trained on the system's specific limitations?

testing approach

  • Sample high-stakes AI decisions: what did human review look like?
  • Review override logs: how often are overrides exercised?
  • Interview reviewers: can they explain what they're reviewing and why?
  • Test the halt mechanism: does it actually work?
  • Check training completion records for operator AI literacy
07 dog
Cybersecurity & AI-Specific Security Risk
adversarial attacks, prompt injection, model extraction, data poisoning

AI systems face attack vectors that traditional cybersecurity frameworks don't fully cover. AI-related security incidents surged 56.4% in 2024–2025. Prompt injection attacks exploited legitimate GenAI tools at 90+ organizations to steal credentials. OWASP Top 10 for LLMs (2025 edition) is the practitioner reference.

attack vectors to assess

  • Adversarial examples: crafted inputs that fool the model
  • Data poisoning: malicious manipulation of training data
  • Model inversion / extraction: extracting training data or IP
  • Prompt injection: hijacking LLM behavior via malicious input
  • Model theft: replicating proprietary models via repeated querying

audit questions

  • Are AI systems in scope for vulnerability management and pen testing?
  • Is red-team testing conducted using adversarial inputs?
  • Are model weights, training pipelines, and inference APIs access-controlled?
  • Are prompt injection controls in place for GenAI systems?
  • Are AI-enabled fraud scenarios in the incident response playbook?
08 dog
Third-Party & Vendor AI Risk
due diligence, contract audit rights, sub-processor chains, embedded SaaS AI

Most organizations use more AI from vendors than they build themselves — but vendor AI often receives far less scrutiny than internal models. Sub-processor chains (vendor uses another AI provider) create hidden risk. AI features embedded in approved SaaS tools (Slack, Salesforce, Microsoft 365) frequently bypass the AI intake process entirely.

ISACA 6-step vendor framework

  • 1. Classify vendors by AI risk tier (I–III)
  • 2. Pre-contract due diligence: data use, retention, opt-out rights
  • 3. Contract review: audit rights, breach notification, deletion obligations
  • 4. Ongoing monitoring: accuracy, drift, bias — not just uptime
  • 5. Incident response integration: vendor AI incidents in your IR process
  • 6. Sub-processor mapping: assess the full chain

key audit questions

  • Are AI vendors classified by risk tier with proportionate scrutiny?
  • Do vendor contracts include audit rights?
  • Are SaaS embedded AI features inventoried and assessed?
  • Is training data used by vendors without authorization?
  • Are model change notifications required contractually?
  • Is vendor concentration risk assessed?
09 dog
Regulatory Compliance
EU AI Act readiness, sector-specific rules, conformity assessment

The regulatory landscape is accelerating. EU AI Act high-risk obligations are live for enforcement August 2, 2026. Prohibited practices have been banned since February 2025. Sector-specific rules (SR 11-7 for financial services, FDA AI guidance for medical devices, EEOC for employment) layer on top of cross-cutting frameworks.

audit questions

  • Is the org tracking EU AI Act obligations by tier?
  • Is there a conformity assessment plan for high-risk AI systems?
  • Are prohibited practices actively monitored for inadvertent use?
  • Is ISO 42001 or NIST AI RMF adoption tracked against a roadmap?
  • Are sector-specific AI regulations tracked (SR 11-7, FDA, EEOC)?
  • Is there a regulatory change management process for AI law?

self-assessment vs. notified body

  • Self-assessment: available for most Annex III systems
  • Notified body required: biometrics, safety-critical product AI
  • Self-assessment is rigorous — not a lighter path
  • Notified body capacity is thin — engage early
  • EU Declaration of Conformity must be signed and supported by a technical file
10 dog
Monitoring, Incident Response & Model Lifecycle
drift detection, IR plans, retraining triggers, decommissioning

AI systems degrade over time as real-world data drifts away from training data. Model drift monitoring is not optional — it's how you detect performance degradation before it becomes a customer harm event or a regulatory incident. The absence of drift monitoring in production is a High finding.

audit questions

  • Is ongoing performance monitoring in place for deployed models?
  • Are alerting thresholds set for accuracy degradation, data drift, bias drift?
  • Is there an AI incident response plan (with defined escalation paths)?
  • Are retraining triggers and approval processes documented?
  • Is there a model decommissioning process?
  • Are post-incident reviews documented?

monitoring components

  • Data drift: are input distributions shifting vs. training data?
  • Model drift: is predictive performance degrading over time?
  • Bias drift: are fairness metrics moving as data composition changes?
  • Infrastructure monitoring: latency, availability, API error rates
  • Output auditing: sampling AI outputs for quality review
11 dog
Agentic AI & GenAI-Specific Risks (Emerging)
autonomous agents, goal constraints, prompt injection, GenAI use policy

Agentic AI (autonomous multi-step AI systems that take actions in the world) is the fastest-evolving audit frontier in 2025. These systems can make sequences of decisions and take real-world actions without per-step human review. The audit challenges are traceability, goal alignment, and guardrail governance. Standard audit techniques don't yet cover this domain adequately.

audit questions

  • Are AI agents governed with defined goal constraints and guardrails?
  • Are agent actions logged with traceability to reconstruct decision chains?
  • Are human-in-the-loop overrides in place for autonomous AI actions?
  • Is prompt injection risk assessed for GenAI systems?
  • Is there a GenAI acceptable use policy preventing data exfiltration?

GenAI-specific risks (NIST AI 600-1)

  • Hallucination / confabulation
  • Prompt injection and jailbreaking
  • Training data memorization and extraction
  • Intellectual property in training data and outputs
  • Information integrity (deepfakes, synthetic content)
  • Human-AI configuration / automation over-reliance

top findings from enterprise AI audits

Critical
No enterprise-wide AI policy or acceptable use policy
83% of organizations affected — IIA / Wolters Kluwer 2024 Pulse Survey
Critical
AI system inventory incomplete — shadow AI projects (marketing GenAI tools, SaaS-embedded AI) running without inventory tracking or approval
Deloitte 2025 Hot Topics #1 governance gap; 56–80% employee shadow AI usage rate
Critical
No AI governance body with defined authority — only 18% of organizations have an enterprise-wide council with actual decision authority over AI
IAPP AI Governance Profession Report 2025
High
No independent model validation before deployment — model developed and validated by the same team, no conceptual soundness challenge
Common in non-financial-services organizations; SR 11-7 requirement for banks
High
Bias testing absent or ad hoc — training data not reviewed for representativeness across demographic groups; no documented remediation process when bias detected
EU AI Act Art. 10 requirement; NYC LL 144, Colorado SB 205, Illinois AI law
High
Explainability and transparency failures — inability to explain model outputs; inadequate logging; adverse action notices legally insufficient
ECOA / FCRA requirement; GDPR Art. 22; EU AI Act Art. 13
High
Human oversight nominal, not substantive — human-in-the-loop procedures exist on paper but operators rubber-stamp AI outputs without meaningful review
EU AI Act Art. 14 material compliance requirement
High
Third-party vendor AI not assessed — no visibility into what external AI vendors log, retain, or share with sub-processors; vendor contracts lack audit rights
ISACA 2025; Deloitte agentic AI guidance
High
Model drift monitoring absent — no production performance monitoring; no alerting thresholds for accuracy degradation or bias drift
ISACA algorithm audit guidance 2024
High
No AI incident response plan — AI-specific attack scenarios not included in IR playbooks; no tabletop exercises run
Widely noted gap; NIST AI RMF MANAGE function requirement
Medium
EU AI Act gap — no conformity assessment plan for high-risk AI systems ahead of August 2, 2026 enforcement deadline
Wolters Kluwer; Trilateral Research 2024
Medium
Internal audit function itself lacks AI skills — 60%+ of CAEs cite skills gap; only 4% report substantial AI progress within the IA function itself
AuditBoard 2025; IIA Vision 2035 survey

EU AI Act compliance timeline

Aug 2024
Act enters into forceEU AI Act officially in effect. Organisations should begin AI inventory and risk tiering immediately.
Feb 2, 2025
Prohibited AI practices banned — NOWSocial scoring by public authorities, subliminal manipulation, real-time remote biometric surveillance in public spaces (with narrow exceptions). Check your AI portfolio now.
Aug 2, 2025
GPAI model obligations applyGeneral-purpose AI model providers (foundation model providers) must comply with transparency and governance codes of practice.
Aug 2, 2026
High-risk AI system enforcement — the main deadlineArticles 9–17 fully mandatory for most Annex III systems. Conformity assessments must be complete. Fines: up to €15M or 3% of global annual turnover. This is the critical milestone for most internal auditors.
Dec 2, 2027
Extended deadlineSome Annex III categories (biometrics, critical infrastructure, education, employment, migration) get extra runway.
Aug 2, 2028
AI in regulated productsAI embedded in regulated products (medical devices, machinery, toys, lifts, etc.) must comply.

key resources — updated june 2025

IIA
IIA AI Auditing Framework (Sept 2024)
The primary practitioner reference. Four-part structure, three lines model, checklist. Download the PDF.
NIST
NIST AI RMF 1.0 + Playbook
72 subcategories across GOVERN / MAP / MEASURE / MANAGE. The playbook has action-level templates per subcategory.
NIST
NIST AI 600-1 — Generative AI Profile (Jul 2024)
200+ suggested actions across 12 GenAI risk categories. Extends AI RMF 1.0 for LLMs and GenAI systems.
EU
EU AI Act — High-Level Summary & Full Text
Article-by-article reference. Key: Arts. 6, 9, 10, 13, 14, 16, 43. Aug 2, 2026 is the enforcement date.
ISACA
Leveraging COBIT for AI Governance (2025)
ISACA white paper mapping COBIT to 5 AI risk categories and 7 governance enablers. Free download.
KPMG
Auditing Artificial Intelligence (KPMG 2025)
Illustrative AI Risk and Controls Guide + Trusted AI Framework covering EU AI Act, ISO 42001, NIST AI RMF.
Deloitte
Internal Audit's Role in Strengthening AI Governance
Deloitte's practitioner guide. Covers Trustworthy AI pillars, adversarial testing, agentic AI governance.
PwC
Assurance for AI (Jun 2025)
First independent AI assurance service under AICPA standards, aligned to NIST AI RMF + ISO 42001.
AIGL
AI RMF 1.0 Controls Checklist (214 pages)
Maps AI RMF 1.0 to 58 auditable compliance controls with indicators, frequency definitions, and templates.
ISACA Journal
A Proposed High-Level Approach to AI Audit (2024)
Practitioner-focused audit approach from ISACA's journal. Free to read for ISACA members.
Wolters Kluwer
How Internal Audit Must Respond to the EU AI Act
Practical guide for IA functions preparing for EU AI Act obligations. Covers checklist, governance, and readiness.
AuditBoard
Closing Internal Audit's AI Gap
Honest assessment of where IA functions are and aren't on AI readiness. Required reading for CAEs.
EU AI Office
GPAI Code of Practice — Final (Jun 2025)
Binding compliance standard for general-purpose AI model providers under the EU AI Act. Covers transparency, copyright, and systemic risk obligations. Compliance creates presumption of conformity with the Act.
OWASP
LLM Top 10 v2.0 (2025)
Updated adversarial risk reference for LLM applications. Adds agentic AI attack vectors, AI supply-chain attacks, and expanded prompt injection guidance. Essential for AI security testing workstreams.
OMB / White House
OMB M-25-21: Federal AI Governance (Apr 2025)
Replaces M-24-10. Sets minimum AI risk management practices, inventory requirements, and governance expectations for federal agencies. Useful enterprise governance benchmark regardless of sector.
IAPP
AI Governance in Practice Report 2025
Survey of 1,000+ privacy and AI governance professionals. Key finding: only 18% of organisations have a governance council with real authority. Benchmark data for governance maturity assessment.
NIST
NIST Agentic AI Safety Guidance (2025)
Supplemental guidance extending AI RMF 1.0 for autonomous and multi-agent AI systems. Covers goal alignment, tool-use governance, and traceability requirements for agentic AI.
Colorado AG
Colorado SB 205 Implementation Guidance (2025)
Algorithmic fairness obligations for high-risk AI in employment, insurance, credit, and housing. Effective February 1, 2026. Colorado AG issuing compliance guidance — monitor if operating in Colorado.
MITRE
MITRE ATLAS — Adversarial ML Threat Matrix (2025)
Updated adversarial threat landscape for AI/ML systems. Covers model poisoning, model extraction, prompt injection, and agentic AI attack chains. Essential reference for AI red-team and security testing.

what's new — key guidance since may 2025

EU AI Act — GPAI obligations live (Aug 2025)

  • General-Purpose AI model providers must now comply with transparency, copyright, and systemic-risk obligations
  • GPAI Code of Practice (final, Jun 2025) is the compliance standard — adherence creates presumption of conformity
  • Next major deadline: high-risk AI system enforcement — August 2, 2026
  • Prohibited AI practices ban (Feb 2025) is active — verify your portfolio is clean

Agentic AI — new audit frontier (2025)

  • NIST AI Safety Institute released agentic AI guidance extending AI RMF 1.0
  • OWASP LLM Top 10 v2.0 adds agentic AI and AI supply-chain attack vectors
  • MITRE ATLAS updated with agentic AI attack chains
  • Standard audit checklists do not yet adequately cover autonomous AI systems

US federal AI governance shift (Jan–Apr 2025)

  • EO 14179 (Jan 2025) rescinded Biden AI EO — agencies revising policies with innovation-first framing
  • OMB M-25-21 (Apr 2025) replaces M-24-10 with updated federal AI governance requirements
  • FFIEC joint AI risk management guidance for financial institutions in development
  • EEOC and DOL updated guidance on AI use in employment decisions

State AI laws — enforcement approaching (2025–2026)

  • Colorado SB 205: algorithmic fairness obligations effective February 1, 2026
  • NYC Local Law 144: employment AI bias audit requirement — enforcement active now
  • 15+ states introduced AI governance bills in 2025 legislative sessions
  • Texas, Illinois, Virginia bills advancing — monitor for applicability to your operations

questions to ask your organisation — internal AI governance audit

Q1 dog
Governance & Strategy
board ownership · AI risk appetite · ethics operationalisation · oversight structures
dog

Dog #1 says: "If governance questions get blank stares, you already have your first finding. The absence of an answer is the answer."

ask your organisation

  • Has the board formally approved an AI risk appetite statement — and can anyone produce it on request?
  • Who is the named executive accountable for AI governance (CAIO, CDO, CRO) — and what authority do they actually have over AI deployment decisions?
  • Does an AI governance committee exist with defined membership, a regular meeting cadence, and documented decision rights?
  • Are AI ethics principles documented AND operationalised — is there evidence of their application in real decisions, not just on a website?
  • Is there a formal intake and approval process for new AI use cases before deployment — not retrospectively after go-live?
  • Has the board or audit committee received a substantive AI risk briefing in the last 12 months?
  • Are all employees aware an AI policy exists, trained on its requirements, and able to articulate what is and isn't permitted?
  • How does the organisation define "AI system" — and is that definition consistently applied across all functions and business units?
Q2 dog
AI Inventory & Shadow AI
completeness · SaaS-embedded AI · intake process · employee usage
dog

Dog #2 says: "Ask for the inventory first. Then go find what's not in it. The gap between those two lists is your audit."

ask your organisation

  • Can you produce a complete, current AI system inventory — right now, not after two weeks of data collection?
  • Does the inventory include AI features embedded in SaaS platforms already in daily use (Microsoft Copilot, Salesforce Einstein, Workday AI, ServiceNow, Slack AI, Zoom AI)?
  • Does it include AI used in pilots, proofs of concept, and team-level experiments — not just production systems?
  • Does it capture AI used by third-party vendors who process your data on your behalf?
  • When was the inventory last actively reviewed for completeness — and by what method?
  • Is there an approved intake process for evaluating new AI tools before employee adoption?
  • Have you conducted a shadow AI discovery exercise (network logs, SaaS subscription audit, expense reports, employee survey) in the last 12 months?
  • What is the consequence for an employee who adopts an unauthorised AI tool — and has it ever been enforced?
Q3 dog
Data Governance & Quality
training data lineage · consent for AI use · bias at source · retention · cross-border transfers
dog

Dog #3 says: "The question nobody thinks to ask: was this personal data collected with consent for AI training specifically — not just 'we had it and needed it'?"

ask your organisation

  • Is training data lineage documented to source — for each production AI system individually, not just in aggregate?
  • Was personal data used for AI training collected under consent that explicitly covers that specific AI/ML use?
  • Are data quality controls applied before data enters AI training or inference pipelines — not just assumed from upstream systems?
  • Are training datasets reviewed for demographic representation and potential bias before use?
  • Are retention schedules specific to AI training data, model weights, and inference logs — distinct from general data retention policy?
  • If a data subject exercises a right of deletion, does that obligation extend to their data in training sets and model outputs?
  • Are cross-border transfers of data used in AI training or inference assessed under GDPR Chapter V or equivalent transfer frameworks?
  • Has a data protection impact assessment (DPIA) been completed for every AI system that processes personal data?
Q4 dog
Model Risk & Validation
independent validation · performance metrics · fairness testing · model drift · model registry
dog

Dog #4 says: "The question that exposes everything: who validated this model before it went live — and were they on the same team that built it?"

ask your organisation

  • Are AI models validated independently before deployment — by someone demonstrably independent of the development team?
  • Is there a documented AI/ML development lifecycle with mandatory approval gates before production deployment?
  • Are performance metrics (accuracy, precision, recall, F1) formally defined, baselined, and monitored in production for each model?
  • Are models tested for fairness across relevant demographic groups — before deployment and on an ongoing basis post-deployment?
  • Are model cards or technical documentation files maintained and current for each production AI system?
  • What triggers a mandatory model revalidation — and is that criteria documented in policy, not left to developer discretion?
  • Is there an AI model registry tracking the version, status, owner, review date, and risk classification of every production model?
  • How would you know if a deployed model's performance had meaningfully degraded — today, right now?
Q5 dog
Regulatory Compliance Readiness
EU AI Act tiers · prohibited practices · sector-specific rules · conformity assessments
dog

Dog #5 says: "August 2, 2026 is not a vague future date — it is 14 months away. No conformity assessment plan means you have a regulatory exposure finding right now, today."

ask your organisation

  • Have all AI systems been formally classified under the EU AI Act risk tiers — prohibited, high-risk, limited, minimal — with documented rationale?
  • Have prohibited AI practices (banned February 2, 2025) been positively verified as absent from the AI portfolio?
  • Is there a conformity assessment plan for each high-risk AI system — with a named owner, timeline, budget, and milestone tracking?
  • Are sector-specific AI regulations tracked and mapped to your AI inventory (SR 11-7 for financial services, FDA AI guidance for medical devices, EEOC guidance for employment AI)?
  • Is there a regulatory change management process to track new AI laws globally and update policies and controls accordingly?
  • For employment-related AI: are mandatory bias audits being conducted per NYC Local Law 144? Has Colorado SB 205 applicability been assessed ahead of the February 2026 effective date?
  • Is an ISO 42001 certification roadmap in place — and if not, is NIST AI RMF alignment being tracked against a maturity target?
  • Does legal counsel have full visibility into the AI system portfolio and its regulatory exposure across all applicable jurisdictions?
Q6 dog
Human Oversight & Automation Bias
substantive review · override rates · halt mechanisms · operator training
dog

Dog #6 says: "Are your human reviewers actually reviewing — or just approving? One of these is compliance. The other is theatre."

ask your organisation

  • For each high-stakes AI decision (credit, hiring, healthcare, benefits): can you describe specifically what human review looks like in practice — not in policy documentation?
  • What is the AI output override rate across high-risk systems — is it tracked and reviewed for automation bias patterns?
  • Have halt and override mechanisms for high-risk AI systems been actually tested — not just designed and documented?
  • Are human reviewers trained on the specific failure modes and limitations of the AI systems they are accountable for reviewing?
  • Has automation bias (reviewers defaulting to AI recommendations without independent scrutiny) been identified and treated as a formal operational risk?
  • Can the organisation produce evidence of substantive human review through logs, recordings, or sampling — not just process documentation?
  • What is the escalation path when a human reviewer disagrees with an AI output — and is it actually used?
  • For EU AI Act-scoped high-risk systems: is Article 14 compliance documented, evidenced, and tested — not just asserted?
Q7 dog
Third-Party & Vendor AI Risk
contract audit rights · sub-processor chains · SaaS AI features · ongoing monitoring
dog

Dog #1 says: "'We bought it from a vendor' is not a control. It's a statement that the work hasn't been done yet."

ask your organisation

  • Is there a vendor AI risk classification process — and does contract review include AI-specific due diligence before signature?
  • Do vendor contracts explicitly include: audit rights, data use restrictions, model change notifications, breach notification obligations, and deletion rights?
  • Do you know who your key AI vendors' sub-processors are — and have those sub-processors been assessed?
  • Are AI features embedded in existing approved SaaS platforms (Copilot, Gemini Workspace, Slack AI) captured in the AI inventory and subject to the same risk assessment as standalone tools?
  • Is vendor AI performance monitored on an ongoing basis — not just assessed at contract signature?
  • Is vendor concentration risk for AI dependencies assessed — what is the business impact if a key AI provider becomes unavailable?
  • When a vendor updates their AI model or changes data processing practices, how does your organisation find out — and is there a contractual notification requirement?
  • Has any vendor used your organisation's data to train their AI models without explicit authorisation?
Q8 dog
Cybersecurity & AI-Enabled Fraud
pen testing scope · prompt injection · IR plans · deepfake and voice cloning risk
dog

Dog #2 says: "AI-enabled fraud is not a future scenario — voice cloning BEC and AI-generated phishing are live threats right now. Has your SOC run a tabletop on this yet?"

ask your organisation

  • Are AI systems explicitly included in the scope of the vulnerability management program and penetration testing — not treated as out-of-scope applications?
  • Is red-team testing (adversarial prompting, boundary testing, jailbreak attempts) conducted for GenAI-facing systems?
  • Is there an AI-specific incident response plan — or are AI security incidents handled through a generic IR plan not designed for AI attack vectors?
  • Have AI-enabled fraud scenarios (deepfake voice/video for BEC, synthetic identity fraud, AI-generated spear phishing) been added to tabletop exercises in the last 12 months?
  • Are model weights, training pipelines, and inference APIs protected with appropriate access controls and monitored for unauthorised access?
  • Are prompt injection controls in place and tested for any GenAI system that processes external user input?
  • Does the SOC have specific playbooks for AI security incidents — model poisoning, LLM data exfiltration, prompt injection attacks?
  • Has AI-enabled fraud risk been quantified, reported to senior leadership, and included in the enterprise risk register?
Q9 dog
GenAI Acceptable Use & Agentic AI
acceptable use policy · hallucination risk · agent guardrails · blast radius
dog

Dog #3 says: "The most dangerous question in AI governance right now: what is the maximum autonomous action your AI can take without human approval — and does leadership know the answer?"

ask your organisation

  • Is there a GenAI acceptable use policy covering: which data classifications are permitted in which tools, prohibited uses (regulated data, client data, credentials), output verification requirements, and policy violation consequences?
  • Are employees trained on GenAI hallucination risk and their personal responsibility to verify AI-generated content before acting on it or distributing it?
  • For AI systems that take autonomous multi-step actions (code execution, emails, API calls, data writes) without per-step human approval: are goal constraints and guardrails formally documented and tested?
  • Are all agentic AI actions logged with sufficient detail to reconstruct full decision chains for incident investigation and audit purposes?
  • What is the largest consequential autonomous action any AI system in your environment can take without human approval — and is that level of autonomy explicitly sanctioned by leadership?
  • Is there a formal process for defining and approving the scope of authority granted to AI agents before deployment?
  • Have LLM-integrated systems been tested for prompt injection, jailbreaking, and data exfiltration risks?
  • Does the organisation have a definition of "agentic AI" — and is there a complete inventory of systems meeting that definition?
Q10 dog
Internal Audit Function Readiness
AI literacy · audit universe coverage · IA's own AI use · audit committee reporting
dog

Dog #4 says: "You're auditing AI while AI is being used to audit. The function that ignores its own AI governance gap is the least credible auditor in the room."

ask your organisation

  • Does the IA function have members with sufficient AI technical literacy to conduct credible AI audits — or is external SME engagement required and budgeted?
  • Is AI risk formally included in the annual risk assessment and IA audit universe — with appropriate risk ratings and audit frequency?
  • Does the IA charter explicitly address AI audit responsibilities, authority, and access rights to AI systems and model documentation?
  • Is the IA function using AI tools in audit work — and if so, are those tools subject to the same governance standards being audited elsewhere in the organisation?
  • Is AI risk formally escalated to the audit committee at least annually — with specific findings, management responses, and remediation tracking?
  • Has the IA function assessed itself against the IIA AI Auditing Framework — and documented the gaps?
  • Is continuous monitoring or data analytics being used to provide ongoing assurance over AI system outputs — beyond point-in-time audits?
  • Is the IA budget, staffing plan, and skills development roadmap adequate for the AI audit workload coming in FY2026 — EU AI Act enforcement, state AI laws, GenAI proliferation, agentic AI expansion?

dog golden rules for auditing artificially intelligent nerds

  1. Start with the inventory. You cannot govern, audit, or regulate what you cannot see. An incomplete AI inventory is both the most common finding and the root cause of most others. Spend at least 30% of your planning phase on shadow AI discovery.
  2. Absence of governance is a Critical finding. No AI policy, no governance body, no ethics principles — these are not "opportunities for improvement." They are the absence of the control environment. Rate them accordingly.
  3. Test human oversight; don't just document it. "Human-in-the-loop" that consists of rubber-stamping is not human oversight. Sample high-stakes decisions. Interview reviewers. Check override logs. The ghost of human review is not the same as human review.
  4. The August 2026 EU AI Act deadline is not hypothetical. Fines up to 3% of global revenue for high-risk AI non-compliance. If your organization has high-risk AI systems affecting EU persons and no conformity assessment plan, that is a regulatory exposure finding today.
  5. Bias audits are now legal compliance, not ethics aspirations. NYC LL 144 is live. Colorado SB 205 is coming. EU AI Act Art. 10 is coming. The demographic breakdown of model outputs is now a legal artifact, not a social good.
  6. Vendor AI gets the same scrutiny as internal AI. "We bought it from a vendor" is not a control. Vendor contracts without audit rights, data use restrictions, and model change notifications are gaps, not just best practices.
  7. Build your AI audit skills before AI audits you. IIA data: GenAI use in internal audit doubled to 40% in one year. You will be auditing AI while AI is being used to audit. Close your own function's AI skills gap before it closes you.
  8. Pat the dogs. Take a break. The finding you keep missing is usually in the process you understand best — precisely because you stopped questioning it.