What Counts as CX ROI Evidence Before Funding an AI Project
A three-lever funding gate for AI projects. Cost-to-serve, capacity-reallocation, and customer-felt functionality are the only places a project can cash out. If the vendor pitch does not disclose evidence for at least one lever with a pre-declared baseline, the request is positioning, not operating proof.
Quick decision summary
Five plain-language checks for a go or hold decision
- What claim are we testing?
- The three-lever rule (cost-to-serve, capacity-reallocation, customer-felt functionality) is a sufficient funding gate for any enterprise AI project request, applied before approval rather than in the post-mortem.
- Who is the named peer?
- PwC-Anthropic alliance (May 14, 2026; Paul Griggs, Andy Crowder on record), Starbucks-NomadGo Automated Counting (deployed Sept 2025, retired May 18-22 2026; Brian Niccol on record), KPMG-Anthropic alliance (May 19, 2026; Bill Thomas, Rema Serafi on record).
- Source strength
- T1 T1 (named buyer on record with primary source)
- Where this may not apply
- The rule is calibrated for regulated-enterprise buyers (financial services, healthcare, professional services). Consumer-product AI projects may have additional levers (engagement, retention) that this framework does not cover. The funding-gate discipline assumes a Director-of-AI-Platform-grade decision authority with budget over a multi-million-dollar program. Pilot-scale decisions under $250K may not need the full gate.
- Recommended decision
- Fund the project only against one named lever with one disclosed metric and a pre-declared baseline. Expand to the second lever only after the first moves measurably (above noise, against the baseline) within two operating quarters. If no lever is named with disclosed evidence, decline funding and request a re-pitch with the gap specified.
A vendor walks into the funding meeting. The deck says AI transformation, agentic acceleration, customer-centric reinvention. The ask is eight figures over two years. The room is supposed to decide yes or no by Friday.
The decision question is not whether the technology works. The decision question is whether the project, if funded, will move something the customer can feel. Everything upstream of that question is interesting. Nothing upstream of it is sufficient.
This article names what counts as evidence before you sign, and what does not.
Every AI project cashes out through one of three levers
For regulated-enterprise funding decisions, I would force every customer-facing AI request to declare one of three cash-out paths:
- Cost-to-serve: the same service outcome delivered at lower dollar cost per unit. Same call resolved, same claim adjudicated, same ticket closed, fewer dollars spent.
- Capacity-reallocation: the same team producing more serving capacity at stable quality. More claims processed, more cases triaged, more accounts opened, with the same headcount or fewer.
- Customer-felt functionality: a product or service capability the customer experiences directly. Faster resolution from their frame, an answer they could not get before, a new opt-in feature that did not exist last quarter.
A project that cannot point to at least one of these three is not an AI project for budget purposes. It is a technology purchase looking for a sponsor.
What evidence looks like per lever
The funding gate is not a slide that says “we expect to reduce cost-to-serve.” Expectation is not evidence. Evidence is something with a number, a baseline, and an audit trail.
Cost-to-serve
Look for: dollar per unit, pre-deployment versus post-deployment, same scope, same accounting basis. The pre number must come from the buyer’s own ledger, not the vendor’s modeled estimate. The post number must be reproducible after the implementation team leaves.
Disqualifying language: “expected savings of,” “industry benchmarks suggest,” “ROI calculator shows,” “early indicators.”
Capacity-reallocation
Look for: hours freed per week multiplied by units of work the freed hours actually produced, plus a queue-aging or backlog delta to confirm the freed capacity is being absorbed by demand and not by slack.
PwC’s May 14 announcement claimed insurance underwriting compressed from 10 weeks to 10 days, with the explicit framing that the compression “opens lines of business that were not previously economically viable.” That is the right shape: a capacity number that the buyer can verify, paired with a demand statement about what the new capacity will be used for.
Disqualifying language: “saves N hours per employee per week” with no statement of what the freed hours produced. Time saved that no one absorbs is not capacity. It is slack.
Customer-felt functionality
Look for: a metric measured from the customer’s frame, not the operator’s. Time-to-resolve from the moment the customer first contacted, not from the moment a routing engine assigned the case. Defect rate as the customer experienced it, not as the QA team scored it. Opt-in for a new feature with retention past 30 days.
Disqualifying language: internal user satisfaction scores, employee NPS for the tool, vendor-issued sentiment analysis, anything measured before the customer touched it.
What is not evidence
The most common failure mode in AI funding pitches is substituting upstream signals for customer-felt outcomes. Watch for these specifically:
- Deployment count. “Rolled out to 276,000 employees” is a logistics achievement. It does not tell you whether one customer experienced a different outcome. KPMG-Anthropic and PwC-Anthropic both lead with workforce-rollout numbers in their May 2026 announcements; both are honest about what they are (Pass 1 inputs), but the funding gate has to be Pass 2 (customer-felt outputs).
- Token-cost reduction. “Sixty-one percent cheaper inference” is a supplier-economics fact. It changes the vendor’s margin. It does not necessarily change the customer’s experience or the buyer’s cost-to-serve, because token cost is rarely the binding constraint in an enterprise deployment.
- Pilot results that have not survived scale. Starbucks deployed NomadGo’s Automated Counting across North America in September 2025 on the strength of pilot signal. Nine months later, in May 2026, the tool was retired because field execution did not match pilot performance. A pilot that worked in a controlled environment is a hypothesis about scale, not evidence at scale.
- Vendor-issued case studies without buyer attestation. If the customer name appears in the announcement but the customer did not publish their own metric, the case is procurement-stage proof (they bought) and not deployment-stage proof (it worked). Andy Crowder’s on-record quote about Advocate Health’s PwC-Anthropic engagement is a procurement signal. The funding gate for the next phase still needs Advocate-issued numbers.
- Efficiency metrics that stop at the operator. “Resolution time reduced by 40 percent” sounds like customer-felt functionality but often is not, because the clock starts when the AI assigned the case, not when the customer first reached out. Read the metric definition before treating the metric as evidence.
The funding gate, stated
Decline funding unless the pitch discloses:
- One named lever out of the three, with a stated reason the project addresses that lever specifically.
- One pre-declared baseline measured from the buyer’s own data before any vendor intervention.
- One target metric with a measurement window (one quarter or two quarters, not “ongoing”), a measurement method, and a named owner who will publish the result.
- A failure condition that, if met, ends the program. “We will continue funding regardless of outcome” is not a project. It is a subsidy.
If those four are present, fund the project against one lever only. Do not approve a multi-lever program on the strength of a single-lever case. Expand into a second lever after the first has moved measurably, above noise, within the declared window. Expand into the third lever after the second has done the same.
If those four are not present, decline and request a re-pitch with the gaps specified. This is the cheapest moment to surface the gap. The cost rises every quarter after sign.
What this gate prevents
It prevents the failure mode that Starbucks ran into between September 2025 and May 2026: a tool deployed at scale on the strength of strategic narrative (“Back to Starbucks turnaround lever”), with no pre-declared cost-to-serve baseline, no capacity-reallocation metric the field operators verified, and no customer-felt functionality signal beyond stockout-prevention assumptions. When the field reality failed to match the pilot, the only honest response was retirement. The cheaper response was a tighter funding gate nine months earlier.
It also prevents the inverse failure mode, which is rejecting a project that has real evidence on one lever because the pitch does not have evidence on all three. Most AI projects start by moving one lever. The discipline is to fund against that one lever, measure it, and expand from there. Demanding full-spectrum CX transformation in the first pitch deck guarantees no project ever clears the gate, which is the same outcome as funding everything that gets pitched: the gate stops working.
Monday morning
For an AI Platform Director with budget authority over a multi-million-dollar program, the practical move this week is to take every active funding request and answer four questions for each:
- Which of the three levers is this project built to move first?
- What is the pre-declared baseline for that lever, sourced from our own ledger?
- What is the target metric, the measurement window, and the named owner?
- What failure condition ends the program?
Requests with crisp answers to all four can proceed to the next gate. Requests without crisp answers go back to the vendor with the specific question they cannot currently answer. That move alone, applied consistently for one quarter, will resolve more funding-decision ambiguity than another round of vendor demos will.
The gate is not adversarial. It is the cheapest form of help a buyer can give a vendor: tell them, before they ask for the check, exactly what evidence will get the check signed.