Introducing Financial Audit Bench Read the post →

Introducing Financial Audit Bench

Financial Audit Bench chart comparing strict pass rates and mean cost per run across eleven AI models
Jerry Huang, Jim Burton, Pranav Pillai, Matt Van Buren, and Alex Wang
FinancialAuditBench: strict pass rate versus mean cost per run for 11 models FinancialAuditBenchPass@1 · mean of 8 runs per task 20% 30% 40% 50% 60% 70% 80% $0 $1 $2 $3 $4 $5 $6 $7 Mean cost per run (USD) Strict pass rate Claude Opus 5.569.2% Claude Fable 5.165.8% Grok 4.647.2% GPT-5.6 Sol44.2% GPT-6 Astra43.6% GPT-5.6 Terra40.3% Muse Spark 1.339.4% DeepSeek V4.1 Flash33.1% GLM-5.229.4% Kimi K328.9% Gemini 3.8 Flash27.5% Closed weights Open weights

TLDR

  • Today, we’re introducing Financial Audit Bench (FAB), an open-source benchmark that tests whether agents can perform long-horizon tasks in financial statement audit. We’re releasing the paper, benchmark code, datasets, and live leaderboard so others can reproduce, challenge, and extend our results.
  • Evaluation on eleven frontier models showed low reliability despite high partial-credit scores. No model strictly passed more than 69% of audit procedures, while average strict-pass rates fell to <10% on judgment-intensive tasks such as testing revenue (8%) and accounts receivable (4%).
  • Financial Audit Bench is the first project from Modus Labs, a team that brings leading PhDs and Applied AI Engineers together with the top minds in public accounting. This includes Jim Burton – ex-Chief Auditor at Grant Thornton – who has helped translate audit-quality principles into concrete standards for how agents should be incorporated in the modern, AI-native accounting firm.

Motivation

We founded Modus to deliver the highest quality audit in a fraction of the time. AI is fundamental to advancing Audit Quality, enabling teams to analyze a higher volume of transactions and interrogate the idiosyncratic risks in each business, while eliminating rote questions and data requests which otherwise burden clients.

While time savings are straightforward to quantify, Audit Quality has historically been an area with less precise measures – ironic, given our industry.

As models advance and more firms deploy agents in audit, FAB examines three practical questions: which tasks are appropriate for agents; which models are best suited to those tasks; and which failure modes should auditors anticipate when reviewing agent-generated work.

What’s in the Benchmark

Task Selection

The tasks in Financial Audit Bench (FAB) are designed to model work typically assigned to junior (staff-level) auditors during the substantive fieldwork phase of a financial statement audit. They map to how work gets done at accounting firms.

How a FinancialAuditBench task mirrors audit fieldwork at a CPA firm FinancialAuditBenchOne task = one workpaper 01 PLAN 02 EVIDENCE 03 FIELDWORK 04 REVIEW At a CPA firm Audit plan A senior sets the risks and the procedures to perform. Client binder Client records, bank and customer confirmations. Staff workpaper A staff auditor performs and documents the work. Senior review A reviewer checks key tie-outs and conclusions. In Financial- AuditBench Instructions A workpaper template and planning context. Synthetic binder ~179 files of planning, client, and third-party evidence. Agent workpaper One bash tool, up to 200 model calls. CPA rubric ~10 checks. Passes only if every check is met.

In this benchmark, we defined tasks spanning 8 core audit areas. Agents were evaluated on their ability to perform end-to-end procedures using a fixed data room of client-provided and third-party evidence; they cannot request additional information during a task.

Procedures per Audit Area
FinancialAuditBench audit areas and the procedure types each workpaper requires FinancialAuditBench8 audit areas · 45 workpapers Tie-out & reconciliation Third-party confirmation Vouching & tracing Recalculation Cutoff Estimates & judgment Presentation & disclosure Cash Accounts receivable Accounts payable Debt Fixed assets Revenue Journal entry completeness Journal entry risk testing

Building the Engagements

In FAB, we built a high-quality dataset of representative audit tasks, validated by CPAs and audit executives. To protect sensitive client records, we generated 6 engagements using a computational graph of ~250 nodes and ~1,000 edges. Every document is derived from the graph, so balances, dates and transactions agree across them. The graph’s starting values came from aggregate statistics over historical audits, released under a (3.65, 10⁻⁵) differential-privacy guarantee at the engagement binder level. Auditors reviewed each round of generated binders for inconsistencies and implausible content. As in real audits, some imperfections were left in; rubrics don’t grade the parts of a workpaper those imperfections affect.

Grading

At firms today, senior auditors spot-check key details in a workpaper rather than re-performing audit procedures. We mirror this process in FAB. Every workpaper has a reference answer prepared by at least one CPA-licensed auditor and reviewed by a senior auditor with more than ten years of experience. Its rubric evaluates ~10 checks per workpaper across key balances, calculations and conclusions. The grader finds each check by table column and row key, not by fixed cell address, so an agent can add or reorder rows without being penalized. Where a table can be completed correctly in more than one way, GPT-5.6 Luna judges the extracted content.

Strict pass requires every applicable check to be met. Partial credit is the share of checks met. Across the eight runs, pass@k measures whether at least one of k attempts passes, and pass^k measures whether all k do.

Sample Tasks

Results

We evaluated eleven frontier models (eight closed-weight, three open-weight) on 45 workpaper completion tasks. Each model attempted every task eight times, producing 3,960 trajectories. Agents received a bash interface as their only tool and were limited to 200 model calls per attempt.

Every model satisfied more than 85% of applicable rubric checks on average, but no model achieved strict pass on more than 69% of task attempts. Strict pass requires satisfying every applicable check in a workpaper.

Model Performance

FinancialAuditBench: strict pass rate versus mean cost per run for 11 models FinancialAuditBenchPass@1 · mean of 8 runs per task 20% 30% 40% 50% 60% 70% 80% $0 $1 $2 $3 $4 $5 $6 $7 Mean cost per run (USD) Strict pass rate Claude Opus 5.569.2% Claude Fable 5.165.8% Grok 4.647.2% GPT-5.6 Sol44.2% GPT-6 Astra43.6% GPT-5.6 Terra40.3% Muse Spark 1.339.4% DeepSeek V4.1 Flash33.1% GLM-5.229.4% Kimi K328.9% Gemini 3.8 Flash27.5% Closed weights Open weights

At time of writing, Claude Opus 5.5 led all models with a 69% strict pass rate, and at $3.30 a run it cost half as much as the runner-up, Claude Fable 5.1 (66%). Thereafter we see a 19-point drop to Grok 4.6 (47%), and the field bottoms out at 28%. The best open-weight model, DeepSeek V4.1 Flash, strictly passed 33% of task runs at $0.14 a run.

Task-Level Performance

FinancialAuditBench strict pass rate by audit area FinancialAuditBenchStrict pass rate · mean of 11 models Journal entry completeness 94.5% Accounts payable 68.2% Cash 56.1% Fixed assets 54.2% Debt 52.1% Journal entry risk testing 16.9% Revenue 7.6% Accounts receivable 4.2% Accounts payable has 3 tasks (manufacturing only); every other area has 6.

Across all eleven models and eight runs per model-task pair, the average strict-pass rate was 94.5% for journal entry completeness, compared with 52.1% to 68.2% across accounts payable, cash, debt and fixed assets. Performance was weakest on procedures estimated to require the most auditor time: journal entry testing (17%), revenue (8%) and accounts receivable (4%). These procedures often require the preparer to distinguish among several plausible sources of evidence – a recurring issue in our qualitative failure review.

Reliability Across Repeated Runs

FinancialAuditBench reliability: pass^k from k=1 to k=8 FinancialAuditBenchpass^k · 45 tasks × 8 runs 0% 20% 40% 60% 80% 1 2 3 4 5 6 7 8 Attempts, all required to pass (k) Tasks passed 9 other models Claude Opus 5.544.4% 69.2% Claude Fable 5.131.1% 65.8%

To quantify dependability of an agent across runs, we use pass^k, which measures success on all k attempts. Model performance was not consistent across runs; Opus passes all eight attempts on only 44% of tasks, Fable on 31%, and Kimi K3 on none.

Common Failure Modes

To understand where agents went wrong and assess whether the rubrics were fair, auditors compared agent-completed workpapers with CPA-reviewed reference workpapers and reperformed selected procedures. Three patterns stood out:

Three recurring agent failures, shown as flagged workpaper excerpts FinancialAuditBenchIllustrative excerpts 01InsufficientretrievalREVENUE · CUTOFF TESTSale testedInvoice dated 12/30Source reviewedShipping cutoff scheduleShip dateNot in binder 02Weighing evidencepoorlyACCOUNTS RECEIVABLETestedBalance at 12/31Vouched toCash receipts reportEvidence sourceClient-prepared 03Completing tasks withinsufficient informationSUBSEQUENT EVENTSPeriod requiredJan–Feb 2026Period reviewedDec 2025ConclusionNo events noted
  1. Insufficient retrieval. Agents analyze only the most obvious source document and conclude, on that basis alone, that the required evidence is unavailable. On a revenue cutoff procedure, for example, an agent reviewed a custom shipping cutoff schedule that lacked the necessary detail, although the full shipping log in the same binder contained it.
  2. Weighing evidence poorly. Agents relied on client-prepared evidence without obtaining the external corroboration required by the procedure. For example, an agent vouched receivables to a client-prepared cash-receipts report but did not agree the reported receipts to bank activity. The report therefore showed what the client recorded, not whether the cash had been received.
  3. Completing tasks with insufficient information. Unlike insufficient retrieval, this failure arises when the information a procedure requires is genuinely absent from the binder. Rather than noting that the procedure cannot be performed as designed, agents document work using whatever information is available. For example, an agent that had not been provided any subsequent-period information documented a subsequent-events procedure using current-year transactions.

Limitations

FAB evaluates staff-level procedures across six synthetic engagements in two industries. It does not capture every dimension of a production audit, including client interaction, requests for additional evidence, supervision and rework, or engagement-level judgment. Its results should therefore be interpreted as evidence about performance on the evaluated tasks – not readiness to conduct an audit autonomously.

Additionally, FAB’s rubrics primarily measure whether selected balances, calculations and conclusions are correct; they do not measure how easily another auditor can review the resulting workpaper. In qualitative review, auditors found that agents often documented their process at length while leaving the relevant evidence, its significance and the resulting conclusion less clear. Sentence fragments and poor cell formatting created additional review friction. These observations suggest that reviewability is a distinct dimension of workpaper quality—one that future evaluations should measure alongside correctness.

Guidance for our Peers

We built Financial Audit Bench to answer practical questions: where can firms use agents today, where do they still require substantial oversight, and how do those answers change as models improve? The results suggest that agents can complete meaningful portions of staff-level audit work, but their outputs still require professional review.

We therefore recommend deploying agents as co-pilots rather than autonomous preparers. Human-in-the-loop systems will likely address the evidence-selection and escalation failures observed in FAB, although the benchmark did not directly evaluate interactive auditor intervention.

The development of this paper also highlighted the velocity at which new model releases are changing the landscape of AI in audit. Within a matter of weeks, the benchmark’s leading strict-pass rate rose from below 50% with Grok 4.6 to 69.2% with Opus 5.5. That pace favors model-agnostic infrastructure that allows firms to evaluate and adopt stronger models and agent frameworks as they emerge.

Appendix: Key Terms

Audit

Engagement
The audit of one company for one reporting period, usually a fiscal year.
Engagement binder
The materials for one engagement: planning documents, client-provided records, third-party evidence and blank workpaper templates.
Workpaper
A spreadsheet, organized by a predefined template, where auditors document the procedures performed, evidence examined and conclusions reached.
Substantive fieldwork
The audit phase in which auditors respond to assessed risks by performing procedures and gathering evidence.
Client-provided (PBC) records
Records, financial documents and ledgers prepared and supplied by the audited company.
Third-party evidence
Evidence sent directly to the auditor by outside parties, such as bank confirmations.

Benchmark

Rubric check
One graded data point, calculation or conclusion. Rubrics average about ten per workpaper.
Strict pass
A workpaper passes only if every applicable check is met.
Partial credit
The share of applicable checks met.
pass@k
Whether at least one of k attempts strictly passes.
pass^k
Whether all k attempts strictly pass.
Differential privacy
A formal guarantee that adding or removing any one historical audit binder barely changes the released statistics. FAB’s guarantee is (ε = 3.65, δ = 10⁻⁵).

Filed under: AI & Automation