Why your firm's AI seats aren't turning into workflows

Separate sculptural forms join into a continuous ribbon, suggesting individual AI use becoming a shared workflow.

A year after a firm buys Claude or ChatGPT seats, the picture is usually similar. One associate has a set of prompts for reading a CIM that saves her an afternoon per deal. A VP uses it to tighten IC memo drafts. A few people tried it once, got a confident wrong answer, and went back to doing the work by hand. Ask how the firm uses AI for first-pass screening and you get four different answers, none of them written down.

Most firms are at this stage, and it's a normal one. The tools work for the people who use them well. The step from personal habit to shared process just hasn't happened yet.

The pattern: individual use that never compounds

Each person prompts differently, so the same CIM produces a different summary depending on who ran it. The best method belongs to whoever found it. When that associate is on vacation, or leaves, the method goes with her, because it lives in her chat history.

The output is also hard to check. One analyst asks for page references and another doesn't. A reviewer looking at a figure in a draft memo often can't tell whether it came from the CIM, the model's arithmetic, or a guess.

Leadership can't see the value either. The seat count and the login report show who opened the tool. They don't show whether deals got screened faster or memos needed fewer corrections.

So individuals get a little faster, and the firm doesn't get better at anything, because nothing anyone learned has been shared.

The numbers say this is normal

Broad individual use with thin firm-level results shows up across the large surveys.

In BCG's AI at Work 2026 survey of 11,749 people, 74% of frontline white-collar employees were regular AI users (daily or several times a week), up from 51% in 2025. Among those regular users, 42% said they save a full day or more each week. But 66% said they get limited or no guidance on what to do with that time, and BCG's own conclusion is that saved time "leaks out of the organization" unless someone tracks it and puts it to use. All of these numbers are self-reported.

At the firm level, the picture is flatter. An NBER working paper, Firm Data on AI (2026), surveyed nearly 6,000 senior executives in the US, UK, Germany, and Australia. 69% of their firms actively use AI, yet nine in ten executives reported no impact on productivity or employment over the past three years. Those are executives' own assessments.

You may have seen the MIT figure that 95% of organizations get zero return from generative AI. The MIT NANDA report behind it is labeled preliminary, draws on 52 interviews and 153 survey responses, and applies the 95% to custom, task-specific tools, where success means a marked and sustained productivity or P&L impact. For general-purpose tools like ChatGPT and Copilot, the report finds wide use (over 80% of organizations have explored or piloted them) and says they improve individual productivity, not P&L. That second finding is the one that fits a firm that bought seats.

Private markets look the same. In an Apex Group survey of 105 senior private credit leaders, 85% said AI is fully embedded in their private credit activities. But only 7% reported true cross-functional ownership of AI, and 49% placed oversight with the CTO or CDO (Apex blog). Apex is a fund administrator and the answers are self-reported. Its own summary was that fewer firms "have redesigned the underlying processes."

Why it stalls

Everyone prompts differently, and the best method stays with one person

OpenAI's State of Enterprise AI 2025 found that the heaviest users (the 95th percentile) send six times as many messages as the median worker. People who used AI across about seven kinds of tasks reported five times more time saved than people who used it for about four. The report comes from the vendor and the savings are self-reported, but the spread matches what we see inside firms: one or two people get most of the benefit, and nobody else can use their methods.

Output is hard to check

In a 2023 field experiment with 758 BCG consultants (HBS summary), people using GPT-4 on tasks within its abilities finished 12.2% more tasks, 25.1% faster, with quality rated over 40% higher. On a task built to fall outside those abilities, they did worse: Ethan Mollick's write-up of the study reports about 84% got it right without AI and 60 to 70% with it. That was one firm, one day, and a 2023 model. The lesson for an investment team is that the same tool helps on one task and hurts on the next, so the output needs a defined check.

Checking takes time, too. In BCG's 2026 survey, 52% agreed that AI had increased the time they spend reviewing and correcting AI output. In Allvue's 2026 GP Outlook, 59% of respondents named accuracy and reliability concerns as a barrier to AI. We cover which underwriting tasks are safe to hand over in where AI still isn't reliable in underwriting.

Leadership can't see the value

In Microsoft's 2024 Work Trend Index, 59% of leaders said they worry about quantifying the productivity gains from AI. In Gallup's AI indicator for US employees, only 25% said their organization has communicated a clear plan for integrating AI. When everyone uses the tool their own way, there is no single process to measure.

Nobody owns it, and few people were trained on real work

Only 36% of respondents in BCG's 2026 survey feel properly trained, unchanged from 2025. In Microsoft's 2024 survey, only 39% of AI users had received AI training from their company. In the Allvue survey, 64% named limited internal expertise as a barrier.

Without an approved, shared method, people fill the gap themselves. Microsoft found that 78% of AI users (not of all employees) bring their own AI tools to work. The MIT report found that only 40% of companies had bought an official LLM subscription, while workers at over 90% of the companies surveyed reported regular use of personal AI tools for work. For a firm handling CIMs and borrower data, that can mean confidential documents going into personal accounts. We cover the risks in is it safe to put a CIM into ChatGPT.

What works: make one workflow standard

The research points one way, though it is correlational. In BCG's 2026 survey, employees at companies that redesign workflows end to end were more likely to see measurable business improvement than those at companies that mainly roll out tools (67% vs 43%). Bain's Global Private Equity Report 2025 says the leading firms tie AI to "a short list of strategic priorities." Neither proves a method, but both point toward changing a few processes properly instead of spreading tools wider.

Here is how we do it. These steps are our method, built from practice; no study has tested this exact sequence.

1. Pick one workflow

Choose weekly work with a template and a clear reviewer: a first-pass deal screen, a diligence request list, or a review of a borrower's compliance certificate. Give it one owner who can make decisions. Our post on building a deal screening memo from a CIM walks through one example.

2. Test it on real past work

Take a handful of completed deals where you already know the right answer. Run the workflow on their source materials and compare the output to what your team actually produced. Look for missed risks, wrong figures, and anything a reviewer would have to fix.

We do this because people are poor judges of their own speed and accuracy with AI. In a METR study of 16 experienced software developers, participants expected AI to speed them up by 24%, believed afterward it had sped them up by 20%, and in fact took 19% longer. That is a small study in another field, but it is reason enough to test against known answers.

3. Build in your templates and review steps

The output should arrive in your firm's memo format, with each figure citing the document and page it came from. Write down who checks what before anything reaches IC. Then save the whole thing as a shared project or custom assistant in the firm's own AI account, so everyone runs the same version. OpenAI reports that about 20% of its enterprise messages already go through Custom GPTs or Projects, and that the most widely used ones capture institutional knowledge.

4. Train the team on that workflow

General AI training has limits. In the BCG consultant study, a short briefing on GPT-4's limits did not help people on the task the model got wrong. Train on the actual workflow, with live sessions on real deals, and record them so new hires can catch up. Get managers involved: Gallup found that employees whose manager strongly supports AI use were 1.7 times as likely to use it a few times a week.

5. Hand it over

The owner gets a written playbook, the recordings, and a short list of things to measure, such as time per deal and how many corrections reviewers make. Re-run the past-work test when the model or the template changes.

What this looked like at one firm

For a private credit manager we worked with, we built a deal screening and diligence request workflow into the firm's own AI account. The team was testing the first working version within three days and sent back structured feedback. Associates then proposed their own additions, which we folded into the shared workflow. Other teams asked for the materials. The investment team had five live, recorded training sessions.

The associates' additions matter most to us. Each one went into the version everyone uses, where before it would have stayed in one person's chat history.

A checklist for this quarter

  • Ask three people how they use AI on the same task, and compare the answers.
  • Pick one workflow that happens weekly and already has a template.
  • Name one owner who can make decisions about it.
  • Pull a handful of completed deals with their source materials.
  • Run the workflow on them and compare the output to what the team produced.
  • Write down the review step: who checks what before anything goes to IC.
  • Record a baseline now for time per deal and reviewer corrections.
  • Confirm the workflow runs in the firm's approved AI account, under your data controls.
  • Schedule live training on that one workflow, and record it.

Still choosing a tool? See Claude vs ChatGPT vs Copilot for investment firms.

Where we come in

This is the work we do: one workflow, tested on your past deals, built into your templates and review steps, with training and a handover. For credit teams, see our AI consulting for private credit, including candidate workflows, pilot scope, and the recent engagement behind the approach.

Sources