AI agents in finance: is your data ready?
AI agents in finance need approved sources, set definitions, limited access and a reviewer. Start with one department's monthly variance explanation.
In this guide
AI agents in finance are software that takes steps in a finance workflow, such as running a query, comparing records or drafting an explanation for approval, as well as answering questions. Your data is ready for one when a single recurring workflow has approved sources, written definitions, access limited to the task and a named finance reviewer.
Take a monthly expense-variance explanation for one department. The agent pulls the approved actuals and budget, runs predefined queries to calculate the variance and drafts the explanation. Finance sets the definitions, decides who sees which fields and approves the explanation before the department head gets it. The agent writes nothing to the ledger or the budget.
When we interviewed finance teams, three organizations described the data groundwork AI depends on as unfinished: figures that differ between accounting and analytics systems, documents that must be normalized as they are ingested, and records still being consolidated into a store that software can query. Those are reported situations, not measured ones.
So don't start by choosing a platform. Start with that one workflow, because approved actuals and budget give you a known answer to test against.
Biggest takeaways
- Start with one recurring explanation. A department's monthly variance has approved inputs and a result you can check; an open request to ask about company finances doesn't.
- Let fixed queries calculate and the model explain. The numbers come from predefined, version-controlled queries. The model drafts the explanation and labels any cause it can't support.
- Limit what the agent can see. Grant only the fields the task needs, and test every place its output lands.
Set up the pilot once
- Name the owners. A finance owner for definitions and acceptance, a technical owner for queries and connectors, an owner after handover, and a reviewer who doesn't configure the agent.
- Set the explanation threshold. Decide which variances need an explanation, as an amount, a percentage of budget or both, and whether either test is enough to trigger one.
- List every input and its definitions. Record each source's owner, access approver and reconciliation rule, and write down the rules the agent must apply.
- Give the agent read-only access to the fields the task needs, under its own service account.
- Measure today's effort, if efficiency is part of the business case.
- Run the monthly steps on past months, then decide whether to expand.
Write the source register and definitions
List every input. In the example, posted actuals come from the general ledger, the budget from the planning system or file, department headcount and payroll totals from payroll or HR, and the department mapping from the cost-center hierarchy. For each, record the owner, who approves access, the entity, the period, how it's refreshed, its approval state and how it reconciles.
Then define what counts. Are actuals posted, preliminary or adjusted? Original budget or approved revised forecast? How are shared costs and intercompany charges handled? An expense can be accrued or paid, capitalized or expensed, so write down the rule for each.
Write down what lives in people's heads. In one organization we interviewed, department budget knowledge sat with one person. Keep effective dates for rates and amendments, including decisions made in email, and keep account and period mappings through an ERP migration or chart-of-accounts redesign. Another organization told us it validates its reports again as systems and report writers change. When the ledger and the approved reporting package disagree, keep both references and ask the owner.
Set what the agent can see and do
Approve each integration separately; approval to use an application may not cover connecting it to another system. Include the agent's service account in user access reviews. Department totals usually cover the task: the agent reads headcount and payroll totals by department, never individual pay. Set a minimum group size and suppress payroll totals below it, because in a department of one or two people the total is someone's salary. In one organization we interviewed, interest in querying finance data came with concern about exposing salary information, so treat this as a design requirement.
Keep it read-only for the pilot. Before you let the agent write anything, define its authorization, the transactions in scope, the activity record and how you recover from a bad write.
Run each month in five steps
The example explains September's operating-expense variance for a made-up marketing department, with a made-up threshold: a variance over $10,000 or over 5% of budget needs an explanation, and either test is enough. All names, amounts and run details are illustrative.
Step 1: Pull the sources and check the data cutoff date
The agent pulls each source in the register and records the query version, extraction time and data cutoff date. If September's actuals aren't posted yet, the run labels them preliminary.
Step 2: Calculate the variance with fixed queries
The variance comes from predefined, version-controlled queries or formulas, not SQL the model writes at run time. In the example, September actuals of $412,000 against a September budget of $380,000 give a $32,000 overspend, 8.4% over and above the threshold.
Step 3: Draft explanations for variances over the threshold
The agent gets a short, fixed instruction:
Explain each department variance above the threshold. Use only the attached query results and department totals. Label each cause supported, contradicted or unconfirmed, with the lines or totals behind it. Classify it as budget phasing, a real overspend or underspend, or a misposting to reclassify. Explain favorable variances too. Don't calculate anything yourself.
Budget phasing means the cost landed in the right month but the budget expected it in another, so that month comes in under budget by the same amount. A real overspend stays, and a misposting needs a reclass. Once a variance crosses the threshold, its causes must add up to it, with any remainder shown as unexplained.
| Line | What the agent produced | Status | Reviewer's decision |
|---|---|---|---|
| Actuals | $412,000 posted to marketing cost centers after the September close; data cutoff October 6, run R-117 | Ties to the approved reporting package | Accept |
| Budget | $380,000, September line of the original annual budget | Matches the approved budget | Accept |
| Variance | $32,000 over budget (8.4%), calculated by query | Verified; above threshold | Accept |
| Cause 1 | Event sponsorship, $25,000: event held and expensed in September, budgeted for October | Supported by GL lines; budget phasing | Accept |
| Cause 2 | “Higher spending from new hires” | Contradicted: department headcount and payroll totals match budget | Reject |
| Remainder | $7,000 of the $32,000 not covered by a supported cause | Unexplained | Send back to the department for detail |
| Access check | Department totals only; no individual pay in the text or attachments | Passed | Accept |
Step 4: Review the package and send back what isn't supported
The package holds the totals, sources, mappings, each cause with its status and any changes since the last run. Put it in your normal acceptance process, so the reviewer can challenge a cause without collecting the data again. One team we interviewed that automates document ingestion still has a person review what it extracts, and that's the right default here too.
In the example, the reviewer accepts the phasing cause, rejects the hiring cause because the department totals contradict it, checks the sponsorship isn't a misposting, and sends the $7,000 back to the marketing team. If no answer comes before the deadline, the explanation records it as unexplained.
If the variance review is also a management review control over financial reporting, expect your auditors to test the completeness and accuracy of the report the agent produces and the precision of the review: the threshold and what the reviewer checks. The person who configures the agent's queries, mappings or instructions shouldn't approve its output. Confirm with the control owner what the reviewer must evidence.
Step 5: Approve, send and keep the record
The department head receives only what the reviewer approved:
September spend was $32,000 over budget. $25,000 is an event sponsorship held in September that the budget placed in October, so October's sponsorship line will come in $25,000 under budget. $7,000 is not yet explained; finance has asked your team for detail.
The record keeps the reviewer, the date, run R-117, the query versions, the data cutoff date and each decision, so your auditor can retrace it.
What goes wrong, and what the reviewer does
The math is right and the explanation isn't. That's Cause 2 in the example. Log extraction, calculation, explanation and access failures separately, since each has a different fix.
A source changes after the result went out. A late entry posts to a reported period, or the budget is revised. Find the affected results, rerun the checks, have the revision reviewed and decide what happens to reports already sent. Ask any supplier to show this with a changed record; retrieving today's file doesn't prove it.
The recipient sees more than they should. Restricted fields can leak through a download, a log or a cached answer. Test every output path, name an owner who can revoke credentials, and change someone's access during a run to see what happens.
Nobody maintains it. Source layouts change, permissions expire and departments reorganize. One practitioner we interviewed expected a growing set of custom workflows to become costly to maintain, though no cost was measured. Whoever inherits the workflow needs its operating knowledge and access to its configuration, tests and logs.
Decide whether to expand
Run the monthly steps on two or three months your team already reviewed. Then try to break it with missing inputs, duplicates, changed records, ambiguous mappings, the wrong period, a failed connector and a request from someone without access.
How do you know it worked?
When it reproduces the reviewed results and the reviewer can decide from the package alone. Agree these checks first; model confidence replaces none of them.
- Every result uses authorized inputs, and its totals reconcile to the approved figures.
- Every variance above the threshold is caught, and the failure cases are flagged.
- Each cause is labeled, and the causes plus any unexplained remainder add up to the variance.
- Recipients see only the fields they're cleared for, in every output path.
- If efficiency is part of the case, preparation plus review takes less effort than the baseline.
When should you stop?
When you can't verify the result. Pause if:
- a definition is missing and no owner will settle it;
- an input has no reconciliation rule;
- a tie-out fails.
Record the owner and the evidence needed to resume. If the gap can be isolated, narrow the scope to what you can verify instead of stopping.
Do you need a new data platform first?
Not for the pilot. Controlled extracts are enough to run past months, so decide on infrastructure from the pilot's results. Live monthly use needs a maintained refresh with a data cutoff date. Broader rollout adds maintained feeds with reconciliation checks, governed definition changes, change detection, regression tests and monitoring.
Questions about AI agents in finance
What are AI agents in finance?
Software that takes steps in a finance workflow, such as running a query or preparing a draft for approval, as well as answering questions. A person reviews its output before anyone relies on it.
What data does an AI agent in finance need?
Approved sources it may read, definitions it can apply, a list of actions it may take, and a review step before anyone uses its output.
What does data governance for AI cover in a finance workflow?
For one workflow: the source register, the definition of each measure, the permission boundary, how changed records are handled, and the owners who keep them current.
Do you need a new data platform first?
Not for a first pilot. Controlled extracts of approved actuals are enough for a supervised trial. If reconciled finance data already sits in a warehouse, point the pilot at it; if it doesn't, don't wait for one.
Check one workflow before you pick infrastructure.
Bring one redacted recurring output, such as a variance explanation, with its sources and how it's reviewed today. We'll go through the readiness checks with you and mark which gaps block a supervised pilot. Your team keeps financial definitions, access approval and acceptance.
Sources and scope
- Published
Airframe Field Research. The finding that the data groundwork is unfinished rests on reported accounts from three organizations, counted by organization rather than by interview. Each other field point in the guide comes from one organization and is described where it's used. Four organizations contributed in all. These are reported situations from confidential interviews, not measured results or a representative survey.
Airframe Market Data. Airframe's market research informed the guide's scope only; no market figures or vendor claims are used.
External Benchmark. NIST SP 800-53 (opens in a new tab) is a security and privacy control catalog you can tailor; it doesn't certify a finance implementation. NIST's voluntary AI Risk Management Framework (opens in a new tab) organizes AI risk work under Govern, Map, Measure and Manage. The checks in this guide apply those ideas; they aren't an official NIST checklist or an assurance opinion.
The worked example is illustrative, and the steps are Airframe recommendations, not a customer case study or a measured outcome.
Related guides in How to AI: Finance: AI for account reconciliation: start with one account; AI in auditing: what a reviewer needs to see; AI for private equity: start with administrator oversight.