Quick answerRun a dedicated conversation-design study when quality and containment metrics look acceptable but a softer signal keeps showing up anyway: customers repeating themselves, rephrasing the same request multiple times, or abandoning a conversation that technically ended in a correct answer. Those are usability problems, not accuracy problems, and a metrics dashboard built around correctness and containment will not surface them on its own. A UX study is also worth triggering after any significant redesign of the conversation flow itself, before a small change to phrasing or turn structure gets rolled out broadly based on intuition alone.
Accuracy metrics and usability are different questions
A containment rate and an accuracy score both answer "did the agent get this right and avoid escalating it," which is a necessary but incomplete picture of whether the conversation actually worked well for the customer. It is entirely possible for an agent to reach the technically correct answer while making the customer repeat themselves twice, rephrase a request because the first phrasing was misread, or sit through an unnecessarily long confirmation sequence before getting there. None of that shows up in containment rate measurement, which counts the outcome, not the path the customer took to reach it. A dedicated usability study exists to look at that path directly, usually through transcript review, moderated sessions, or both.
The signal to watch for is friction that survives a correct outcome
The clearest trigger for a study is a pattern where the metrics say the agent is doing fine, but a softer signal keeps recurring: customers rephrasing the same request, asking the agent to repeat itself, or exiting a conversation that ended correctly but took noticeably longer than it should have. This overlaps with, but is a narrower slice of, mid-conversation abandonment rate, which flags conversations customers leave before completion. A usability study should also look at conversations customers stayed in but were visibly unhappy with, since those never show up in an abandonment metric at all but represent the same underlying design problem playing out with a different ending.
From the team
We build production AI systems for startups.
LLM pipelines, RAG, and agent workflows that hold up under real traffic — not just in the demo.
Trigger a study before, not just after, a conversation-flow redesign
The second reliable trigger is proactive rather than reactive: any time your team plans a meaningful change to how a conversation is structured, whether that is a new confirmation pattern, a new way of asking clarifying questions, or a change to how the agent hands off to a human, run a small usability study before the broad rollout rather than relying on team intuition about what will feel natural to a customer. This is the same caution that a CSAT and NPS impact measurement program already applies after the fact; running the equivalent check before a redesign ships catches problems while they are still cheap to fix, rather than discovering them in a satisfaction score weeks later.
FAQ
How often should a UX study run if nothing specific is triggering it? A baseline cadence, once or twice a year, is reasonable even without a specific trigger, since conversation patterns and customer expectations shift gradually in ways a single trigger event will not always catch.
Who should conduct the study, the product team or an outside group? Either can work, but whoever conducts it needs to be someone other than the person who designed the conversation flow being studied, since designers are the worst-positioned people to notice their own design's friction points.
Does a small company need this as much as a large one? The trigger conditions matter more than company size. A small company running a single, stable conversation flow with no complaints may not need a study yet; a small company mid-redesign needs one just as much as a large one does.

