By the Foxnut team · Updated
Statistics
Financial statements, and what producing one costs
What a set of financial statements costs to produce, what regulators find wrong with them, and where machine tagging of the figures breaks: published numbers, with sources and dates.
By the Foxnut team · Updated
What this page records
The question is whether a machine can produce a set of financial statements, and the published evidence splits it cleanly: extracting the figures is largely solved, and deciding what each figure is has not been. On the full US GAAP taxonomy of 17,688 concepts, the best model tested pulls a number and its context out of a filing at an F1 of 0.6932 and then maps it to the correct accounting concept 18.89% of the time. That gap is the whole process. Producing statements is expensive - the reporting regulator estimates 1,695.2 hours and $334,438.81 of cost per US annual filing - but the expensive part is judgement that is audited, published and, when wrong, corrected in public. Each record below carries the figure, who was counted, over what period, and the publication it was read from, so any single row can be checked without trusting this page. What a build in this area has to hand over at the end is what a reporting system’s owner needs to keep it honest.
1,695.2 hours
Estimated burden per Form 10-K response - the annual filing that carries a US registrant's audited financial statements - across roughly 6,740 responses a year, for a total annual reporting burden of 11,425,648 hours
$334,438.81
Estimated cost burden per Form 10-K response, for a total annual cost burden of $2,254,117,579 across the same 6,740 responses
222
Annual reports and accounts reviewed by the UK reporting regulator in its 2024/25 cycle, down from 243 in 2023/24 and 263 in 2022/23, of which 38% were FTSE 350 companies
37%
Share of those reviews that produced a substantive letter asking the company for further information or explanation: 26% within the FTSE 350 and 44% outside it, against 47% overall the previous year
18
Companies required to refer to the review in their next annual report, which the regulator describes as typical where a company restates comparative information in the primary financial statements. One of the 18 affected profit, against five of 26 the previous year
10%
Share of reviews raising impairment of assets, the most frequently raised issue for the third consecutive year, ahead of cash flow statements and financial instruments at 9% each, then presentation of financial statements and revenue at 5% each
0.1889
Best reported accuracy at mapping an extracted financial figure to the correct concept in the full 17,688-concept US GAAP taxonomy, against an F1 of 0.6932 for extracting the figure in the first place. Both are DeepSeek-V3, the strongest model on each subtask
81%
Share of questions that GPT-4-Turbo with a retrieval system answered incorrectly or declined to answer, on a benchmark of 10,231 questions about publicly traded companies with answers and evidence strings attached
95% to 100%
False-positive rate on clean statements in nine of fourteen complete model runs asked whether a set of financial statements is numerically consistent. One run reported 0% false positives; on a rounded variant of the same set that model's recall was 79.0%, against 100% on the unrounded one
| Number | What it measures | Period | Source |
|---|---|---|---|
| 1,695.2 hours | Estimated burden per Form 10-K response - the annual filing that carries a US registrant's audited financial statements - across roughly 6,740 responses a year, for a total annual reporting burden of 11,425,648 hours | Estimate published 31 July 2026 | US Securities and Exchange Commission, Federal Register, 31 July 2026 |
| $334,438.81 | Estimated cost burden per Form 10-K response, for a total annual cost burden of $2,254,117,579 across the same 6,740 responses | Estimate published 31 July 2026 | US Securities and Exchange Commission, Federal Register, 31 July 2026 |
| 222 | Annual reports and accounts reviewed by the UK reporting regulator in its 2024/25 cycle, down from 243 in 2023/24 and 263 in 2022/23, of which 38% were FTSE 350 companies | 2024/25 monitoring cycle, published September 2025 | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| 37% | Share of those reviews that produced a substantive letter asking the company for further information or explanation: 26% within the FTSE 350 and 44% outside it, against 47% overall the previous year | 2024/25 monitoring cycle | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| 18 | Companies required to refer to the review in their next annual report, which the regulator describes as typical where a company restates comparative information in the primary financial statements. One of the 18 affected profit, against five of 26 the previous year | 2024/25 monitoring cycle | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| 10% | Share of reviews raising impairment of assets, the most frequently raised issue for the third consecutive year, ahead of cash flow statements and financial instruments at 9% each, then presentation of financial statements and revenue at 5% each | 2024/25 cycle, with 2023/24 comparatives | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| 0.1889 | Best reported accuracy at mapping an extracted financial figure to the correct concept in the full 17,688-concept US GAAP taxonomy, against an F1 of 0.6932 for extracting the figure in the first place. Both are DeepSeek-V3, the strongest model on each subtask | Submitted 27 May 2025, last revised 17 May 2026 | Wang and others, FinTagging, arXiv:2505.20650 |
| 81% | Share of questions that GPT-4-Turbo with a retrieval system answered incorrectly or declined to answer, on a benchmark of 10,231 questions about publicly traded companies with answers and evidence strings attached | Submitted 20 November 2023 | Islam and others, FinanceBench, arXiv:2311.11944 |
| 95% to 100% | False-positive rate on clean statements in nine of fourteen complete model runs asked whether a set of financial statements is numerically consistent. One run reported 0% false positives; on a rounded variant of the same set that model's recall was 79.0%, against 100% on the unrounded one | Submitted 28 May 2026 | Panda, FinVerBench, arXiv:2605.29586 |
What the process profile says
Six structural dimensions decide whether a process is worth automating, and financial reporting reads at an extreme on three of them: the regulatory exposure is the highest in this family, the errors that matter are concentrated in judgement rather than in volume, and the software market is mature at the layer that formats the output while being demonstrably immature at the layer that classifies it.
| Dimension | What it reads on financial reporting | Source |
|---|---|---|
| Exception variance | Concentrated in judgement, not in volume. The issues that draw a regulator's substantive question are impairment of assets (10% of reviews), cash flow statements and financial instruments (9% each), then presentation of financial statements and revenue at 5% each. Routine postings are not what goes wrong; estimates and classifications are, and they are the part of the work that changes every period. | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| Volume | Set by the calendar, not by a queue. A registrant files once a year, at an estimated 1,695.2 burden hours and $334,438.81 of cost per response across roughly 6,740 filers. Payback is therefore a question about a recurring cycle of known size and a fixed deadline, not about throughput - which is the opposite of the invoice-style processes in this family. | US Securities and Exchange Commission, Federal Register, 31 July 2026 |
| Cost of an error | Paid in public, by the preparer, a year later. A material error becomes a restatement of comparative information in the primary statements plus a required reference to the regulator's review in the next annual report: 18 such references in the 2024/25 cycle, one of which affected profit, against 26 the year before. | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| Reversibility | Reversible on the page, not in the record. The correction mechanism exists and is used, but it runs through the following year's accounts and is legible to every reader of them. Automated consistency checking is not yet the safety net it sounds like: nine of fourteen model runs judging clean statements flagged 95% to 100% of them as erroneous. | Panda, FinVerBench, arXiv:2605.29586 |
| Regulatory exposure | The highest in this family, and it is statutory rather than contractual. The statements must give a true and fair view under section 393 of the Companies Act 2006 and IAS 1.15, they are audited, and a national regulator reviews a sample each year: 222 reviews in 2024/25, 37% of which produced a substantive letter to the company. | Financial Reporting Council, Annual Review of Corporate Reporting 2024/25 |
| Vendor market maturity | Mature at the filing layer, immature at the judgement layer. The output is a mandated machine-readable artefact tagged against a taxonomy of 17,688 US GAAP concepts, and the tooling around tagging and filing is long established. Best reported accuracy at choosing the right concept for an extracted figure is 0.1889, against an F1 of 0.6932 for extracting it. | Wang and others, FinTagging, arXiv:2505.20650 |
Two of those rows carry the argument. The vendor-maturity row says the market is mature about the wrong half of the problem: formatting, tagging and filing a statement is a solved commodity, and classifying a transaction correctly is where the published accuracy collapses. The exception-variance row says the same thing from the regulator’s side - the issues raised most often are impairment, financial instruments and revenue, all of which are estimates and classification decisions rather than arithmetic.
The volume row is where financial reporting parts company with every other process in this family. An invoice queue rewards automation in proportion to how many invoices arrive. A reporting cycle arrives twelve or four or one times a year, at a size known in advance, against a filing deadline. That makes the payback question a labour question about a known peak, not a throughput question, and it is why the honest first measurement is how many hours the close and the reporting pack actually consume rather than how many entries pass through them.
Budget variance, and why it is not the same process
Budget variance reporting looks like the same work - the same ledger, the same period, the same finance team - and it sits on the opposite side of every dimension above. Nothing in a variance pack is filed, audited or attested. There is no taxonomy to map an account to and no external reader who can be misled, so a misclassification costs an internal argument rather than a restatement. Reversibility is immediate: a wrong variance explanation is corrected in the next meeting, not in next year’s comparatives.
The consequence is that the two halves of what people call financial reporting have almost inverted automation profiles. The statutory half is constrained by classification accuracy and regulatory exposure, which is exactly what the benchmark evidence says machines are worst at. The management half is constrained by whether anyone trusts the explanation attached to the number, which is a different problem and a considerably softer one. A studio being asked to automate financial reporting should establish which half is meant before quoting anything, because a variance-commentary tool and a statutory reporting system share a data source and share almost nothing else.
It is also worth saying plainly that budget variance has no measurable search demand at all in the harvested data, in any question frame. That is a signal about how the work is discussed rather than about whether it matters, and it is one reason the section sits here rather than on a page of its own.
What most claims about this get wrong
The genre’s central error is treating financial statement analysis and financial statement production as one capability. Reading a set of accounts and forming a view is a research task with a tolerant error profile. Producing a set is a classification task with an audited one, and the two have opposite failure costs. The FinanceBench result belongs to the first category and is still sobering: on questions about public company filings, with the evidence available to a retrieval system, the strongest configuration tested answered 81% of them wrongly or declined to answer.
The second error is repeating figures whose source has stopped standing behind them. The single most circulated academic claim in this area - that a large language model outperforms human analysts at predicting earnings direction from anonymised statements - comes from a paper that was withdrawn on 20 February 2025, after a co-author identified inconsistencies in the data and analyses while attempting to replicate earlier work. The paper is in the sources below for exactly that reason, and none of its numbers appear on this page. The general rule this page applies is stricter than checking whether a claim is popular: a figure that cannot be traced to a live primary publication with a date is not a statistic, however often it is quoted, and the widely repeated project-failure percentages in this space fail that test entirely.
The third error is reading a low benchmark score as a permanent verdict. The FinTagging concept-alignment number will move, and the extraction number already has. What is less likely to move quickly is the asymmetry between them, because extraction is a text problem with abundant training signal and concept alignment is a judgement problem against a 17,688-item taxonomy where the right answer often depends on facts that are not in the document at all.
What this page cannot show you
It cannot tell any particular organisation what its own reporting cycle costs. The Form 10-K figures are a regulator’s estimate for a filing that most companies never make, and they are quoted here as the only per-statement cost figure published by a named body with a date, not as a benchmark for a private company’s close. Nothing here is accounting advice, and no position is taken on how any transaction should be recognised, measured or presented - that is the standard-setter’s territory and the auditor’s, and a research page has no standing in it.
It also cannot answer whether a given finance function is ready to automate any of this, which turns on whether the underlying ledger data is complete, documented and consistently coded. That is a data-quality question rather than a reporting one, and the wider readiness assessment sits alongside it; both are separate from the question this page answers.
How these figures were compiled
Every record was read from its primary source on 6 August 2026: the Securities and Exchange Commission’s Paperwork Reduction Act notice for Form 10-K, read in the Federal Register text hosted by govinfo; the Financial Reporting Council’s Annual Review of Corporate Reporting 2024/25, read page by page in the report PDF; and three arXiv papers, read on arXiv. The Federal Register figures were read twice under different prompts to confirm the per-response and total values, and the notice does not split the cost burden into internal and external labour, so no such split is claimed here.
Three sources were dropped rather than approximated, because they could not be reached in session: APQC’s published close-cycle benchmarks, which would have given a days-to-close figure; the US Bureau of Labor Statistics Occupational Outlook Handbook entry for accountants and auditors, which would have given a labour-cost figure; and the Commission’s own hosted PDF of Form 10-K, whose burden block was read from the Federal Register notice instead. All three returned server-side refusals. No figure from any of them appears here in paraphrase, and none was taken from a search result summarising them.
One source is cited as a caution rather than as evidence: the withdrawn 2024 paper on financial statement analysis. Two of the three benchmark papers are peer-reviewed venue submissions and one, FinVerBench, is a single-author preprint; that is stated in its record rather than left for a reader to discover.
The verdict
The evidence supports automating the production of financial statements at the extraction and assembly layer, and does not yet support automating the classification layer, which is where the value and the risk both sit. A system that pulls figures from source records, assembles the pack, checks it against the previous period and formats it for filing is doing work that machines demonstrably do well, and it removes a real and recurring cost - the regulator’s own estimate is 1,695.2 hours per annual filing. A system that decides which account a transaction belongs in, or whether an asset is impaired, is operating in the band where the best published accuracy is 0.1889 and where the regulator’s most frequent challenges land.
The reason is structural rather than temporary. Three of the six dimensions above point the same way: the exceptions are judgement calls rather than volume, the errors are corrected in public a year later, and the exposure is statutory. A process with that profile does not reward removing the reviewer; it rewards giving the reviewer less to do and a clearer record of what was decided. That is also what makes the audit trail the design centre of any such build - not the model, and not the interface.
For a finance team weighing this, the useful first question is not which product to buy. It is which half of the work is actually being automated, how many hours the last close consumed and where they went, and whether the ledger data is coded consistently enough for a machine to classify from at all. A team that can answer those three is in a position to buy on evidence. A team that cannot is being asked to automate a judgement process it has not yet measured, and measuring it first is cheaper than any of the alternatives. Tell us what your last close cost and we will help you find where the hours went.
Sources
- US Securities and Exchange Commission, 'Agency Information Collection Activities; Submission for OMB Review; Comment Request; Extension: Exchange Act Form 10-K', Federal Register document 2026-15557, published 31 July 2026 - read on govinfo.gov Retrieved
- Financial Reporting Council, 'Annual Review of Corporate Reporting 2024/25', September 2025 - monitoring statistics at pages 8 to 9, reporting-requirement framing at page 7 Retrieved
- Wang, Qian, Peng and others, 'FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information', arXiv:2505.20650, submitted 27 May 2025 and last revised 17 May 2026 - results tables read in the HTML version of v5 Retrieved
- Islam, Kannappan, Kiela, Qian, Scherrer and Vidgen, 'FinanceBench: A New Benchmark for Financial Question Answering', arXiv:2311.11944, submitted 20 November 2023 Retrieved
- Panda, 'FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification', arXiv:2605.29586, submitted 28 May 2026 - a preprint, not peer reviewed Retrieved
- Kim, Muhn and Nikolaev, 'Financial Statement Analysis with Large Language Models', arXiv:2407.17866, submitted 25 July 2024 and WITHDRAWN 20 February 2025 after a co-author identified inconsistencies in the data and analyses Retrieved
Foxnut Studios works on briefs like this one from Bengaluru and Paris. If you want the shape of that before you talk to anyone, here is what an AI engagement costs before you ask.