Where AI still isn't reliable in underwriting, and what to do

Give a current model a long CIM and it will hand back a readable summary of the business in minutes, with page references if you ask for them. Give the same model a multi-tab workbook and ask it to add up revenue by segment across eight funds, and the published error rates jump. Both results come from the same tool.
Most of the confusion about AI in underwriting comes from treating those as one capability. This post sets out where the 2026 evidence says AI is reliable, where it needs checks, and where we keep it out of scope entirely, with a working pattern for each.
What the 2026 benchmarks measure
Benchmarks are imperfect, and the scores rise every few months. They are still the best public evidence of where models break, so here are the ones closest to underwriting work.
Reading and simple lookups are strong
FinSheet-Bench (March 2026) asks models questions about synthetic spreadsheets modeled on private equity fund structures. On simple lookups, such as finding one value in one row, accuracy was 89.1% across all models tested and 93.6% for the top three. The best model overall, Gemini 3.1 Pro, scored 82.4%, which the authors describe as about one error per six questions.
Models also do well on some analyst work over public filings. On the Vals AI Finance Agent v2 leaderboard (updated 7 October 2026), which tests 927 entry-level analyst tasks on public filings, the best score in the earnings analysis category was 84.8%.
Those numbers support using AI to read, summarize, and answer questions about documents, as long as someone can check the answer against the source.
Aggregation and modeling are weak
The same FinSheet-Bench study shows what happens when the question requires combining numbers. Accuracy on complex aggregation questions was 19.6% across all models and 33.3% for the top three. On the largest file in the set (152 companies across 8 funds), average accuracy across all models fell to 48.6%.
SpreadsheetBench 2 (June 2026) uses a stricter rule: a task only counts as correct if every cell matches the answer file. That is the standard a lender's model has to meet. The best model, Claude Opus 4.6, completed 34.89% of tasks overall and 34.00% of financial modeling tasks. At the cell level, the same model got 89.69% of financial modeling cells right. A workbook that is about 90% correct is still a wrong workbook, and the authors list insufficient inspection of the sheet and selecting the wrong target cell as the main causes of failure.
On Vals Finance Agent v2, the top overall score was 65.40% (Gemini 4 Argon, partial credit scoring), and the best score in the financial modeling category was 34.5%.
Two caveats. These tests use early 2026 models, and newer ones will score higher. The FinSheet-Bench data is synthetic, and the spreadsheets were converted to text. Even so, the gap between reading and modeling shows up in every benchmark above.
Why the errors are hard to catch
A wrong number in a credit memo is a problem. A wrong number that looks right is a bigger one, and that is the usual failure.
The errors are plausible. When ICAEW tested Excel's COPILOT function in 2025, a five-year projection looked reasonable, but one cell read 22,517 where it should have been 22,497, and the final year had further errors.
The output can change between runs. A study of LLMs set to deterministic settings found accuracy varied by up to 15% across repeated runs of the same task. Those were 2024 models, but the lesson holds: one clean run doesn't prove the next one will be clean.
The tone doesn't change when the answer is wrong. NIST's Generative AI Profile names this risk confabulation, meaning confidently stated false content, and notes that people act on it partly because the answer sounds sure.
Similar line items are easy to confuse. We found no study that measures this directly, so treat it as our inference. Both SpreadsheetBench 2 and FinSheet-Bench report row and column misidentification as a common failure. In a deal model, that looks like pulling EBITDA when you meant Adjusted EBITDA, or run-rate figures when you meant pro forma.
What the vendors say
The companies selling these tools are direct about the limits. Anthropic's guide to reducing hallucinations says its techniques reduce errors without eliminating them, and closes with: "Always validate critical information."
The Claude for Excel documentation lists audit-critical calculations without independent verification as a use to avoid. It also says to use the tool only with trusted spreadsheets, because external templates and vendor files can carry prompt injections. A sponsor's model is an external file.
Microsoft's Copilot in Excel FAQ tells users to avoid Copilot for decisions in sensitive areas, finance included, and to verify anything it creates. Its guidance for the COPILOT function, as summarized by ICAEW, was to use native formulas for arithmetic and lookups. Microsoft retired that function on 14 September 2026.
Reliable today, needs checks, keep out of scope
Here is how we sort underwriting tasks when we scope a workflow.
| Reliable today | Needs checks | Keep out of scope |
|---|---|---|
| Summarizing a CIM or data room into your memo template | Extracting figures from CIMs, QoEs, and borrower financials | Populating your model from a sponsor's model |
| Answering questions about a document with page citations | Comparing reported results to budget, prior periods, and covenant levels | Final numbers that go to IC without a person checking the source |
| Drafting diligence question lists from past lists | Spreading financial statements into a template | Credit conclusions and investment recommendations |
| First drafts of investor letters and DDQ responses | Totals, rankings, and anything that counts or sorts | Any calculation where checking costs as much as doing |
The middle column is where most of the value sits, and it is where the patterns below matter.
Patterns that work
Make it quote the source
For every figure or claim, the model returns the exact text it came from and the page number. If it can't find a supporting quote, it says so and leaves the field blank. Anthropic recommends this approach, along with explicitly allowing the model to say it doesn't know. A reviewer checking a quote against a page is fast. A reviewer re-reading a 200-page CIM to find where a number came from is not.
We cover the deal-screening version of this in building a deal screening memo from a CIM.
Put the math in code
Let the model read and let code calculate. The FinSheet-Bench authors recommend exactly this split: work out the structure of the sheet, extract rows one company at a time, then run every calculation in deterministic code. Claude can run Python in a sandbox for this, and in Excel the same idea means plain formulas. Keep the extracted table and the code visible so a reviewer can see what ran.
Tie out, then check what tie-outs miss
Build the standard checks into the workflow. Subtotals sum to totals, the balance sheet balances, the same figure agrees across tabs, and units and signs are consistent (thousands against millions, negatives in parentheses).
A passing tie-out only proves the numbers agree with each other. A wrong input can still balance, especially when the error sits in a plug or an offsetting line. Someone still checks the cells that drive the decision.
Run it twice
Run the same extraction more than once and compare. Anthropic suggests this as a hallucination check, since disagreement between runs is a sign the answer is shaky. Fields that disagree go to a person. Fields that agree still get spot-checked.
Design the review
"A person reviews it" only works as a control if the review is built to catch errors. People defer to automated suggestions. NIST lists this automation bias as a risk, and a 2023 radiology study found readers' accuracy fell sharply when a simulated AI suggested wrong answers, including among very experienced readers. That study was in medicine, and applying it to underwriting review is our own reading.
In practice, the reviewer checks a sample of figures against the source every time, always checks the handful of cells that drive the credit decision, and logs what they find so you can see error rates by field. If checking a task takes as long as doing it, the task is a poor fit for AI.
The task we keep out of scope
Populating a firm's own model from a sponsor's model looks like an obvious use for AI. It is mostly mapping and arithmetic across many tabs, with dozens of similar line items defined differently by each sponsor.
No public benchmark tests this exact task. Our view comes from the adjacent evidence: 34.89% of spreadsheet tasks fully correct in SpreadsheetBench 2, 19.6% on complex aggregation in FinSheet-Bench, a best score of 34.5% in Vals' financial modeling category, and Anthropic's own warning about untrusted external spreadsheets. One wrong mapping can flow into every downstream number, and finding it means rebuilding the work.
In a recent engagement, we kept populating the firm's financial model from a sponsor's model out of scope and treated it as a separate project that needs a higher level of accuracy. If a firm wants it automated, the safer routes are mapping tables and deterministic code, or AI proposing a mapping that code validates and an analyst approves.
Governance in plain terms
Bank model-risk guidance (SR 11-7) was replaced in April 2026. The OCC's release on the new guidance says generative and agentic AI models fall outside its scope, and the agencies plan to issue a request for information on model risk management with attention to AI. Private credit and PE managers were never bound by SR 11-7 directly. Either way, the governance for AI work is yours to write.
The SEC's 2026 examination priorities say examiners will check that firms' statements about their AI capabilities are accurate and that firms supervise their use of AI. "Our AI underwrites deals" is the kind of claim that invites that review.
Write down which tasks use AI, who reviews each output, what they check, and where the prompts and outputs are logged. That record answers most of what an LP, auditor, or examiner will ask. If data handling is the open question, start with whether it's safe to put a CIM into ChatGPT.
How we scope this work
We build these workflows inside the Claude or ChatGPT workspace a firm already uses, with source citations, checks in code, and a named reviewer on every output. Then we train the team and hand it over. AI drafts. Your team decides.
If you want to see how we decide what goes in each column for your firm, read how we scope this work.
Sources
- FinSheet-Bench (arXiv 2603.07316, March 2026)
- SpreadsheetBench 2 (arXiv 2606.29955, June 2026)
- Vals AI Finance Agent v2 leaderboard
- ICAEW, Introducing Excel's new COPILOT function
- Atil et al., Non-Determinism of "Deterministic" LLM Settings
- NIST AI 600-1, Generative AI Profile
- Anthropic, Reduce hallucinations
- Anthropic, Code execution tool
- Claude for Excel documentation
- Microsoft, Copilot in Excel FAQ
- Microsoft, COPILOT function
- RSNA, AI bias may impair accuracy (2023)
- OCC News Release 2026-29
- SEC Division of Examinations, 2026 priorities