Quick answerDecide whether to retire an underperforming AI agent capability using three specific data points reviewed together, not any one alone: sustained low usage relative to the cost of maintaining it, an accuracy or satisfaction score for that specific capability that has not improved despite genuine investment, and a clear-eyed estimate of what retiring it actually saves once you count ongoing maintenance, monitoring, and support overhead, not just the original build cost. A capability that fails on one dimension but not the others usually deserves more investment first, not retirement. Only when a capability is weak on usage, quality, and cost-to-maintain simultaneously does retirement clearly outperform continued investment.
The decision, then the communication
How to retire a customer-facing feature without a backlash is entirely about the external playbook once retirement is already decided: how to communicate it, how to migrate affected customers, how to time the announcement. That post deliberately does not cover how you arrive at the decision in the first place, and that gap is exactly what this framework fills. Getting the decision criteria wrong produces a communication problem no messaging plan can fully solve, since customers can tell the difference between a capability retired for a good reason, explained honestly, and one retired to quietly bury a mistake.
The three-part test
Usage relative to maintenance cost. Low absolute usage is not disqualifying on its own if the capability is genuinely cheap to keep running and serves a small but important segment well. What matters is usage relative to the ongoing cost of keeping it accurate and monitored, since a capability that requires constant tuning to serve a shrinking user base is a worse trade than a capability that quietly works with minimal upkeep for the same user count.
Quality trend despite genuine investment. Distinguish a capability that is underperforming because it has never been properly invested in from one that has received real investment and still has not improved. The first case is an argument for investment, not retirement. The second, a capability where the team has tried reasonable fixes over a meaningful period and the production quality metrics still show it lagging, is a much stronger retirement signal, since it suggests a structural limitation rather than a fixable gap.
True cost of continued operation, not just the original build cost. Sunk cost from building the capability should not factor into the retirement decision at all. What should factor in is the ongoing, forward-looking cost: engineering time spent maintaining it, the support burden from its failure cases, and the opportunity cost of the team's attention being split across it and higher-performing capabilities. A capability that looked worth building at the time can still be worth retiring now if its ongoing maintenance cost has grown faster than its usage.
Building this into a recurring review, not a one-off audit
The framework works best applied on a fixed cadence, reviewing every capability against the same three criteria rather than waiting for a champion to raise the question about one specific feature, which tends to introduce bias toward whichever capabilities happen to have a vocal internal advocate or detractor. Pair this review with whatever ongoing ROI tracking cadence the company already runs, so the retirement decision uses the same trusted numbers leadership already sees elsewhere, rather than a separate one-off data pull that invites disagreement about methodology.
FAQ
What if a capability fails the test but has one very vocal customer champion? Weigh that customer's actual strategic value and contract size explicitly and separately from the general usage data, rather than letting a single loud voice silently override the aggregate signal. If that customer's business justifies keeping the capability alive specifically for them, make that a deliberate, costed decision, not a default outcome of ignoring the framework.
How long counts as genuine investment before declaring a quality plateau? There is no universal number, but a reasonable minimum is at least one full iteration cycle, long enough to ship a meaningful fix, measure its impact, and confirm the trend, rather than judging a capability's ceiling based on its first, unoptimized version.
Should this framework apply to capabilities required for compliance even if usage is low? No. A capability required for regulatory or contractual compliance is not a candidate for this usage-and-ROI framework regardless of how the numbers look, since the retirement decision there is governed by the compliance requirement itself, not by performance data.

