AI ROI That a PE Board Will Accept: Margin Evidence, Not Productivity Claims
A PE board accepts AI ROI that is banked, attributable and repeatable. Hours saved, tool usage and adoption rates don't count until they show up as margin, cash or cost taken out of the budget. Most mid-market AI pilots report the first kind of number and get asked for the second. This guide sets out an evidence standard you can agree with your CFO before you build anything, and a one-page 90-day board pack to report against it.
Why do most AI ROI claims fail the CFO test?
Because they measure activity, and a CFO is looking for outcomes.
BCG's June 2026 analysis of how CIOs and CFOs can prove the value of technology describes the gap well. Finance wants outcomes that are attributable, timely and repeatable. Technology teams worry that waiting for perfect proof costs competitive ground. Both positions are rational, and they only reconcile when the evidence standard is agreed up front.
The scale of the problem is large. BCG's August 2026 piece, Look Past Productivity to Get Real Value from AI, reports that while 82% of CEOs are more optimistic about AI ROI, only about 6% of companies see meaningful value measured in lower costs or higher revenue. Its diagnosis is specific. Productivity gains stay theoretical unless they're tied to a concrete action, such as reducing external provider spend, increasing throughput or removing capacity.
Earlier research points the same way. MIT NANDA's 2025 study found that 95% of the generative AI pilots it analysed delivered no measurable P&L impact. That figure needs careful reading. Many of those pilots had no documented baseline, so even genuine value couldn't be shown. The lesson isn't that AI doesn't work. It's that unmeasured AI can't be defended.
What counts as evidence a PE board will accept?
Sort every benefit into one of three tiers, and never add the tiers together into one headline number.
| Tier | What it is | Examples | Counts toward headline ROI? |
|---|---|---|---|
| 1. Banked | Cash or cost visible in the ledger or budget | Retailer deductions credited back; promotion funding cancelled before launch; external provider spend reduced | Yes |
| 2. Measured and attributed | A KPI moved against a baseline, with a clear causal link, and converted to £ | Forecast accuracy up, and stock write-offs down by a stated amount | Yes, once converted and signed off |
| 3. Soft or enabling | Real but not yet convertible | Hours saved, adoption, faster decisions, lower risk | No. Report separately |
This borrows a discipline BCG describes in its June piece: hard-dollar benefits, like headcount or spend reductions, are categorised separately from soft-dollar ones, and business unit leaders sign off the value themselves. Its bankability test is also worth adopting. Approve strategic spend only when there's an agreed route to banking the gain, such as a budget step-down or a headcount shift.
What's the difference between a claim and a proof?
Here are common claims next to what it takes for a board to accept them. The right-hand column is illustrative.
| The claim | Why a board pushes back | What would be accepted |
|---|---|---|
| "The tool saves 400 hours a month" | Hours saved aren't cash unless capacity is released or spend falls | Named external spend reduced, or a role not backfilled, with the £ figure |
| "Forecast accuracy improved 8 points" | Improvement isn't attributed, and isn't in pounds | Baseline accuracy, post-change accuracy, and the stock write-off reduction it produced |
| "80% of the team uses it weekly" | Usage isn't outcome | Outcome KPI moved, with usage as a supporting indicator |
| "We identified £2M of opportunity" | Identified isn't realised | £ credited or cash recovered to date, and £ still pending kept separate |
How do you set the baseline before you build anything?
Set it in the first week, in writing, and get finance to co-sign it. A usable baseline has four parts:
- The metric, defined precisely (for example, "retailer deductions written off unchallenged").
- The source, meaning the system of record finance already trusts.
- The starting number, usually a trailing 12 months, adjusted for seasonality.
- The review date, and who signs off the result.
This is why we run diagnosis before build. Our AI FlightCheck™ diagnostic is a fixed-price, 2–4 week assessment that ends in a 90-day action plan, and the baseline is part of that work. If you're choosing a partner, ask them to name the baseline before they name the build. Our guide to choosing between strategy and implementation partners covers that question and seven others.
What does a 90-day board pack look like?
One page. The numbers below are illustrative, for a £150M UK CPG manufacturer, and aren't from a client.
| Line | Value | Tier |
|---|---|---|
| Baseline: retailer deductions written off unchallenged (trailing 12 months) | £2.1M | Baseline |
| Credited by retailers to date, confirmed in the ledger | £310K |
|
| Disputed and awaiting credit | £190K | Pending, not counted |
| Analyst hours released | ~180 hours/month |
|
| Headline ROI figure | £310K banked |
Reporting £310K when £500K is technically in play looks conservative. It also makes the number very hard to challenge, which is what earns the next budget round. A board that has seen you exclude the pending £190K will trust the next figure you show them.
Where does this work in UK mid-market CPG and logistics?
The best candidates are use cases where the baseline already sits in a finance or commercial system, and the outcome lands as cash or avoided spend. Two examples from our work qualify:
- Deduction recovery. A £40M UK food brand recovered 60% of previously unchallenged deductions in eight weeks. The evidence is in the ledger. Our guide to AI deduction recovery for UK FMCG covers the mechanics.
- Pre-launch promotion evaluation. A £25M CPG brand saved £180K by cancelling two promotions before launch, which is spend that never left the budget. The method is in our post on killing a losing promotion before it launches, and the wider approach is in AI pricing and promotion optimisation for mid-market UK CPG.
[CONFIRM WITH DELIVERY TEAM before publishing. Replace with one verified example: the baseline figure, how the outcome was measured, who signed it off and the date. For the deduction case, state what "unchallenged deductions" totalled at the start and how the 60% was verified. Do not publish this placeholder.]
Do usage and adoption metrics still matter?
Yes, as leading indicators. They tell you whether an outcome is likely to arrive. They can't stand in for the outcome. If adoption is high and the outcome KPI isn't moving, the pilot is changing behaviour without changing the work, which is a common pattern in the pilots BCG describes as paying off when they aren't.
For a deeper look at building the internal case, see our guide to presenting AI ROI to your board. If you're planning deployment across a portfolio, our 100-day ROI playbook for PE portfolio companies covers the rollout. This post is the measurement standard that sits underneath both.
How do you start?
Agree the evidence standard with your CFO first, then pick one use case whose baseline you can already pull. The AI FlightPath™ Sprint takes a scoped use case to production inside ten weeks, which fits a 90-day proof window. Where you need ongoing ownership of ROI reporting, a fractional Chief AI Officer retainer holds it without a permanent hire.
FAQ
How do you create a board-ready AI ROI narrative?
Start from the evidence tiers: banked, measured and attributed, and soft. Lead with the banked figure, show the baseline it's measured against, list what's pending and excluded, and state the next review date. A narrative built this way is short and hard to argue with.
What should a PE-backed company look for to make sure an AI pilot shows measurable ROI within 90 days?
Pick a use case where the baseline already sits in finance or commercial data you trust, such as deductions or promotion spend. Agree the metric and starting number before the build. Choose a partner who prices against a production milestone and names the baseline in the proposal.
How do you show a PE board that AI spend is tied to margin improvement, not just productivity?
Convert every benefit into one of three tiers and report only banked and attributed outcomes in the headline figure. Show productivity gains separately, and link each one to a specific action, such as reduced external spend or released capacity, if you want it to count.
Is time saved a valid AI ROI metric?
Only once it's converted. Time saved becomes ROI when it removes cost, avoids hiring, or lets you handle more volume without adding capacity, and the conversion is visible in the budget. Until then it's a soft benefit.
What baseline do you need before starting an AI project?
A defined metric, a system of record finance trusts, a starting number over a representative period (usually trailing 12 months), and a named person who signs off the result.
Want to know what your first 90-day board pack would contain?
Start with the AI FlightCheck™ diagnostic: fixed price, 2–4 weeks, ending in a 90-day action plan and an agreed baseline.
