AI in auditing: what a reviewer needs to see
AI in auditing can retrieve evidence and draft control test results. Start with one control; a qualified auditor reviews and decides each conclusion.
In this guide
AI in auditing means handing parts of an audit procedure to AI: it retrieves the evidence, compares it with what the control requires and drafts the test result. A qualified auditor approves the procedure and decides what the evidence supports. The aim is to move retrieval and routine checks to the agent so auditors spend their time on judgment.
Start with one control your team already tests, and rerun it on periods you've reviewed, where you know the right answer. That shows where the agent fails before anyone relies on it. Expand only when it passes the checks below.
When we talked to one internal-audit team, they reported working this way: agents read each control, pull the evidence it calls for and compare it, and an auditor starts every run by hand and checks the output. They planned to schedule runs but hadn't authorized that yet.
Biggest takeaways
- Pick a control with reachable evidence and a known history. It needs agreed criteria, a qualified reviewer and past examples of it working and failing, such as a monthly reconciliation.
- Let the agent retrieve and draft; the auditor concludes. A matching balance says nothing about whether anyone reviewed the reconciliation, so have the agent report each attribute separately.
- Request missing evidence before you call it an exception, and test whatever comes back.
Plan the pilot in six stages
- Pick the control and time today's test. Record how long the manual test takes and what review catches.
- Name an independent reviewer. Choose a qualified auditor who didn't design or run the process under test.
- Set the guardrails. Give the agent read-only access, and keep the evidence in a read-only store with a change log so the agent can't alter its own test evidence. Confirm your data policy allows this evidence to go to the model, and control who can change the agent's prompts, model and connectors.
- Rerun past periods against expected answers. Pick reviewed periods with missing, conflicting, stale and ambiguous support, and write down the right result for each first. Each rerun follows the five steps below.
- Run live tests one at a time, with the same five steps. An auditor starts each run and reviews every result.
- Expand or stop. Expand only when every check under "How do you know it worked?" passes, and stop on any repeated failure.
Each test run, step by step
The example tests a made-up control on one operating account. The March cash reconciliation is prepared within 10 business days of month end, every reconciling item over $5,000 is explained and supported, and someone other than the preparer reviews it within 15 business days. Names, dates and amounts are illustrative.
Run step 1: Approve the procedure and the sample
Write down the account, period, source reports, evidence for each attribute, deadlines, required reviewer and what the test can conclude. Recalculation alone shows neither design nor operating effectiveness.
The agent can draft the evidence list from the control narrative; an auditor approves it, the sampling and any extra procedures before the run.
Select the sample your methodology requires for a monthly control; the example follows one selected month.
Run step 2: Retrieve the evidence and check it against the request
The agent pulls the support for each attribute: bank statement, trial balance, reconciliation workbook, reconciling-item support and review record. For each file it logs the report parameters, version, retrieval time and the page, sheet or cell it used.
Check what arrived against your request index. A "received" flag doesn't prove the right document arrived, and anything the agent can't read stays unresolved.
Run step 3: Run exact checks in code, and let the model read
Ties, totals and dates have exact answers, so run them as code. Use the model for reading: finding the review evidence, matching each explanation to its support and drafting the result. Every attribute gets its own conclusion and names its evidence.
| Attribute | Evidence found | Draft result | How it was checked | Auditor's next step |
|---|---|---|---|---|
| Bank balance agrees with the bank statement | March bank statement and workbook cell C8: $1,301,955 in both | Agrees | Code: exact comparison | Accept |
| Reconciliation ties to the ledger | $1,301,955 less outstanding checks of $23,745 plus the $6,100 deposit in transit = $1,284,310, the March trial balance | Agrees | Code: recalculation and comparison | Accept |
| Prepared within 10 business days | Preparer sign-off dated April 9 | Within the deadline | Code: date comparison | Accept |
| Items over $5,000 explained and supported | Check 1182, $14,200, and a $6,100 deposit in transit: both cleared on the April bank statement. Check 1190, $9,545: explanation only | Exception proposed for check 1190 | Model: matched each explanation to its support | Request support |
| Reviewed by someone other than the preparer | No review record in the retrieved files | Review evidence not found (unresolved) | Retrieval | Request the review record |
Run step 4: Request what's missing, then test what comes back
Check 1190 has no support, and there's no review record. The draft flags both without calling either a deficiency.
The auditor asks for both. Evidence produced after a request needs more scrutiny: test the review against the sign-off timestamp in the reconciliation system or workflow log, not the date written on the record, and check what the reviewer actually examined.
In the example, the workflow log shows the sign-off on April 16, inside the 15-day window. But the reviewer signed off with a query on check 1190 still open, so the review didn't operate at the control's $5,000 threshold. No support arrives for the check, so the auditor records exceptions on both the item and the review attribute and evaluates them.
Run step 5: Review the draft and sign off
The auditor reads the draft beside its evidence and calculations, then accepts it, sends it back or revises the conclusion with a recorded reason.
The workpaper keeps the agent's exact output, the evidence versions, and the model and prompt version. That record is what lets someone retrace the work, as IIA Standard 14.6 expects; a model won't necessarily give the same answer twice. Record who reviewed, when and which version they accepted.
What goes wrong, and what the reviewer does
The tracker says received, but the file is wrong. The internal-audit team we talked to ran into this: their evidence tracker didn't always match what was in the files. Open the support itself before testing.
Two sources agree because they come from one place. If two reports share an upstream source, their agreement isn't corroboration. Record the dependence.
The draft is fluent and wrong. NIST AI 600-1 (opens in a new tab) names confident errors and over-reliance on AI output as risks. Check the evidence behind every result you accept; fluency and confidence scores aren't assurance.
The reviewer built the agent or the process under test. Assign someone else. IIA Standard 2.2 presumes objectivity is impaired when an internal auditor gives assurance over an activity they were responsible for within the previous 12 months. Building the testing agent is a similar self-review risk, and the team we talked to raised the same concern about auditing a process they had designed and run.
Decide whether to expand
How do you know it worked?
It worked when the agent matched your expected answers and preparation, review and rework took less time than the manual test, measured the same way. Agree the bar with the audit lead first, and check:
- Each expected answer came back as written, or the difference is explained. Missing review evidence came back unresolved, and conflicting versions held the conclusion.
- Each workpaper keeps the exact output, evidence versions and model and prompt version, so a competent person can retrace it.
- The drafts the auditor accepted contain no claims their evidence doesn't support.
When should you stop?
Go back to the manual test when any of these keeps happening. A single occurrence means holding that result until the auditor accepts the fix or a documented limitation.
- Support is missing or versions conflict, and nobody can resolve it.
- The agent can't show which evidence supports a result.
- A result can't be retraced from the workpaper.
Add any schedule last, and keep automatic deficiency ratings and autonomous final reporting out of the first pilot.
Improve, build or buy?
- If your audit platform already handles requests, evidence links, approvals and retention, improve it first.
- If your controls and sources are specific and someone can maintain connectors and tests, build a narrow workflow.
- If you buy, validate the tool's methodology, traceability, exception handling and review boundaries.
Airframe's market research describes AI audit agents as covering evidence gathering, control testing and workpaper documentation, for audit firms and internal audit teams alike. Whatever you pick, test it on this control with missing and conflicting evidence and a reviewer who disagrees.
Questions about AI in auditing
How is AI used in auditing?
Use it for the reading-heavy parts of a procedure: finding evidence, matching support to a control requirement and drafting the test result. Exact checks such as totals and dates belong in code, and a qualified auditor decides each conclusion.
How is AI used in internal audit?
Internal-audit teams can use agents to test controls. The agent retrieves the evidence, runs the checks and drafts the result; an auditor starts each run and decides.
What standards apply when AI supports an audit?
For internal auditors, the IIA's Global Internal Audit Standards. For external auditors, PCAOB AS 1105, AS 2605 and AS 2201 in PCAOB audits, and AICPA AU-C 500 and AU-C 610 in other audits.
The IIA's Standard 14.1 (opens in a new tab) requires internal auditors to gather relevant, reliable and sufficient information. If relevant evidence can't be obtained, they must determine whether to identify that as a finding. Standard 14.6 (opens in a new tab) requires documentation that would let an informed, prudent internal auditor repeat the work and derive the same engagement results; internal auditors and the engagement supervisor review it, and the chief audit executive reviews and approves it. Standard 12.3 covers supervision, which the chief audit executive may delegate to qualified people while keeping ultimate responsibility.
PCAOB AS 1105 (opens in a new tab) applies when an external auditor uses information the company produced as audit evidence. The auditor tests its accuracy and completeness, or the controls over them, and evaluates whether it is precise and detailed enough. Whether the external auditor uses internal audit's work at all is a separate judgment under AS 2605 (opens in a new tab) and, in an integrated audit, AS 2201. If your AI-assisted work might feed an external audit, agree evidence expectations with the external auditor early.
For audits not under PCAOB rules, AICPA AU-C 500, Audit Evidence (opens in a new tab), and AU-C 610, Using the Work of Internal Auditors (opens in a new tab), play the corresponding roles.
Where does AI fit in accounting and auditing?
In accounting, AI helps prepare work such as reconciliations, which AI for account reconciliation covers. In auditing, it tests that work; keep the two separate.
Pick the first control to test with AI.
Bring one redacted control narrative, its evidence requirements and a past test with a known gap. We'll work through what an agent could retrieve, what it should flag and how your auditors would review it. Your audit team keeps methodology, judgments and sign-off.
Sources and scope
- Published
Airframe Field Research. The agent workflow, the tracker mismatch and the self-review concern come from one internal-audit team at one organization, in a confidential interview. It is reported practice, not a measured result.
Airframe Market Data. The improve, build or buy section draws on Airframe's market research on AI audit agents.
External Benchmark. IIA Global Internal Audit Standards (opens in a new tab) 2.2, 12.3, 14.1 and 14.6; PCAOB AS 1105 (opens in a new tab), AS 2605 (opens in a new tab) and AS 2201, for PCAOB audits; AICPA AU-C 500 (opens in a new tab) and AU-C 610 (opens in a new tab), for other audits; and NIST AI 600-1 (opens in a new tab), a voluntary risk profile, not an audit requirement.
The worked example is illustrative. The steps are our recommendations, not an audit opinion or a claim that any system meets a regulatory requirement.
Related guides in How to AI: Finance: AI for account reconciliation: start with one account; AI agents in finance: is your data ready?; AI for private equity: start with administrator oversight.