Quick answerRecount the savings case using four categories most headline numbers leave out: ongoing retraining and prompt updates as the product and model landscape change, monitoring and evaluation tooling that runs continuously rather than once at launch, incident response time whenever something goes wrong publicly or quietly, and the human oversight hours that did not disappear but simply moved from doing the work to reviewing the agent's work. A savings claim built only on headcount replaced against subscription cost will almost always look better than the version that includes these four categories, and the gap between the two versions is usually the difference between a marketing number and a decision-grade one.
The four categories that go missing from a first-pass calculation
A first-pass savings calculation usually compares agent subscription or build cost against the fully loaded cost of the headcount it replaces, which understates the true picture in a predictable way. Ongoing retraining and prompt maintenance do not stop after launch, since customer language, product surface area, and model versions all keep changing underneath the agent. Monitoring and evaluation tooling has a real, continuous cost, not a one-time setup cost, because a system you stop watching is a system whose quality you no longer actually know. Incident response time, even for minor incidents that never become public, pulls engineering and support hours away from other work. And human oversight, the review, escalation handling, and quality sampling that a well-run program keeps running, is real labor that a naive before-and-after headcount comparison quietly erases rather than reallocates.
Retraining and monitoring are recurring costs, not launch costs
The most common place a TCO calculation breaks down is treating evaluation and monitoring as part of the build budget rather than a permanent line item. This connects directly to what should actually be budgeted for AI agent maintenance after launch, which makes the case that maintenance is not optional cleanup work but a required ongoing cost center. A savings claim that amortizes the build cost over three years but does not carry a comparable ongoing monitoring and retraining cost across those same three years is comparing a one-time number against a recurring one, which will always favor the agent on paper regardless of whether it is actually cheaper in practice.
From the team
We build production AI systems for startups.
LLM pipelines, RAG, and agent workflows that hold up under real traffic — not just in the demo.
Incident response and hallucination cost belong in the same ledger
Incident response deserves its own line rather than being folded into a vague "support overhead" bucket, because its cost is lumpy: months of nothing, then a cluster of expensive hours resolving something that went wrong in production. This is closely related to, but broader than, translating hallucination rate into an actual dollar cost for the CFO, which quantifies one specific failure mode. A full TCO reality check needs that hallucination cost as one input, alongside the cost of other incident types, and alongside the ongoing human oversight hours, before comparing the total against the original savings claim. Whenever a claim survives that fuller ledger and still shows real savings, treat it the same way you would any other capital decision, using the same rigor as when you first calculated ROI before greenlighting the project, now applied retrospectively rather than as a forecast.
FAQ
Who should own recalculating this once the agent is in production? Whoever owns the original ROI case, typically finance in partnership with whoever runs the platform team, should revisit it on a fixed schedule, not just when someone happens to ask. An annual recount at minimum keeps the number honest as usage and maintenance patterns change.
What if the recalculated number shows the agent is no longer saving money? Treat that as useful information, not a failure to hide. It usually means the scope of what the agent handles should change, or that a specific cost category, often monitoring or oversight, has grown past what the original design accounted for.
Should marketing claims from a vendor be trusted for this calculation? Use them only as a starting point for the categories to check, not as the numbers themselves. A vendor's savings case is built to sell the product, not to survive an internal audit, and it will rarely include your specific retraining and oversight costs.

