Skip to content
Back to blog
July 16, 2026·Stephan Moerman, CEO·2 min read

What a 14-day AI agent audit should deliver

Field Notesaudit pilotgap analysisremediation

“We need AI governance” is too broad to fund, assign, or finish.

A fixed-scope audit makes the problem concrete. In two weeks, a team should be able to establish what agents exist, where the highest-consequence gaps are, and which controls to implement first.

The deliverable is not a compliance badge. It is a verified baseline and a prioritized remediation plan.

What the audit should answer

By the end, leadership should be able to answer:

  • Which agents and agentic workflows are active?
  • Who owns each one?
  • Which models, data sources, tools, repositories, and cloud systems can they reach?
  • What actions can they take without human review?
  • Which policies are enforced at runtime, and which exist only on paper?
  • What evidence is captured per run?
  • Where are cost, reliability, and incident signals visible?
  • Which gaps create the most immediate business, security, or compliance risk?

If the audit produces only a maturity score, it has not answered enough.

Days 1–3: define scope and collect signals

Agree on business units, environments, vendors, and systems in scope. Identify the executive sponsor, operating owners, and technical contacts. Collect existing policies, architecture, vendor records, model usage, service identities, workflow definitions, and available telemetry.

Document what is excluded. A bounded audit is credible when its limits are explicit.

Days 4–7: build and verify the inventory

Combine interviews with technical and commercial evidence. Record each agent's purpose, owner, authority, data, tools, model dependencies, environment, review points, evidence location, and cost owner.

Mark the source and confidence of every record. Verify the highest-risk agents against runtime configuration or logs rather than relying only on a questionnaire.

The inventory will contain unknowns. Treat them as findings, not reasons to delay the report.

Days 8–10: trace representative runs

Select consequential workflows and reconstruct real or safely simulated runs. Follow the request through identity, context, policy, tool calls, approvals, external effects, cost, and outcome.

This exposes the gap between written control and operational evidence. A policy may require review while the runtime has no blocking checkpoint. A log may capture model output but not the downstream action. An owner may exist in a spreadsheet but have no way to suspend the agent.

Days 11–12: rank findings

Prioritize by consequence, likelihood, exposure, reversibility, and evidence quality. Separate immediate containment from longer-term maturity work.

A useful finding includes:

  • the observed condition and evidence;
  • the affected agents and systems;
  • the plausible consequence;
  • the existing control, if any;
  • the recommended action;
  • an accountable owner and target sequence.

Avoid false precision. Use clear risk bands and explain the reasoning behind them.

Days 13–14: agree the control plan

Turn findings into a 30-, 60-, and 90-day plan. Typical first controls include reducing access, assigning owners, adding approval gates, centralizing run evidence, setting budgets, creating incident routes, and retiring unused agents.

The final pack should contain a scope statement, verified inventory, access and dependency map, evidence-gap analysis, risk-ranked findings, and remediation roadmap. It should also list every unresolved assumption and unavailable source.

The value of the audit is not the number of issues found. It is the shared operating truth it creates—and the order in which the organization can improve it.

Know what every agent did — and why.

Start with a 14-day audit of your agents, access, evidence, costs, and human checkpoints.

Explore the audit pilot