Technology and AI

How to Tell if an AI Agent Vendor's Case Studies and ROI Numbers Are Real

Pratik Chothani

Pratik Chothani

Software Development Engineer

·

July 26, 2026

·

4 min read

·

Updated July 26, 2026

How to Tell if an AI Agent Vendor's Case Studies and ROI Numbers Are Real

Quick answer

A vendor case study is real evidence only if it names the specific metric, the baseline it's measured against, the time window, and a customer you can actually talk to — anything vaguer than that is marketing copy. The fastest way to separate real ROI from a polished story is to ask for a reference call with the customer named in the study and ask them directly what changed, what didn't, and what the vendor's numbers leave out.

Why case studies default to best-case, not typical-case

A vendor publishes the deployment that worked, not the median deployment. That's not necessarily dishonest — it's the same selection bias behind every marketing document — but it means a case study answers "is this possible" rather than "is this likely for us." Before treating a number as a forecast, assume it's closer to a ceiling than an average, and ask the vendor directly what percentage of their customers hit numbers in that range.

The four things a real case study names

Vague case studies say "reduced support costs" or "improved efficiency." Real ones name: the specific metric (containment rate, average handle time, cost per resolved ticket), the baseline it's compared to (what the number was before), the measurement window (30 days post-launch vs. 12 months of steady state are very different claims), and the customer's name and industry. If any of these four is missing, ask for it before treating the number as evidence — a vendor unwilling to name a metric precisely usually can't, because the real number is less flattering.

Ask for the reference call, not just the write-up

The single highest-signal step is a direct conversation with the customer in the case study, without the vendor on the call. Ask them three things: what the number actually was before and after, how long it took to get there, and what they'd tell a peer considering the same vendor that isn't in the published version. Vendors who resist connecting you to a reference, or who insist on joining the call, are worth treating with more skepticism than the case study itself warrants.

Check whether the metric would move the business, or just look good

A 40% reduction in "average handle time" sounds impressive, but if handle time was never the bottleneck — if the real cost was in multi-agent architecture hand-offs or escalation volume — it's a vanity metric. Map the vendor's headline number back to your own ROI model before you're impressed by it: does moving that number actually change a cost or revenue line you care about, or does it just look good in a deck.

Ask what didn't work in the same deployment

Every real deployment has a rough patch — a use case the agent couldn't handle, a rollout delay, a metric that moved less than hoped. A vendor who can describe that candidly, in the same case study or on the reference call, is showing you they measure honestly. One who presents an unqualified success story either had an unusually easy deployment or isn't telling you the full picture, and you should assume the latter until proven otherwise.

Pilot the claim before you believe it at your scale

The most reliable check isn't a better case study — it's a scoped pilot against your own baseline, measured the same way you'd measure any agent's ROI, with your own data and your own edge cases, before a full commitment. Vendor numbers are a hypothesis about what's achievable, not a guarantee transferable to your environment; treat the case study as a reason to test, not a reason to skip testing.

FAQ

Should I ask a vendor for their full customer list, not just featured case studies? Yes — a vendor confident in typical results will usually share a broader reference list or aggregate stats across their customer base, not just the two or three flagship stories on their website.

Is a percentage improvement more trustworthy than an absolute number? Neither is inherently more trustworthy; both need the baseline. A "60% improvement" from an unusually bad baseline is a different claim than 60% from an already-decent one, so always ask what the starting number was.

How many reference calls should I do before trusting a vendor's ROI claims? Two or three, ideally in a similar industry or use case to yours, is usually enough to tell whether the published case study is representative or an outlier.

Do case studies from AI vendors tend to overstate results more than typical B2B software? Not inherently, but agentic AI outcomes are harder to measure consistently than simpler software metrics, which gives more room for selective framing — so the verification bar should be at least as high, not lower.

Accelate walks every prospective customer through our own reference customers directly, with no one from our team on the call, because a claim you can't verify isn't a claim worth acting on.

Related posts