A product owner tells you the Claude assistant 'feels slow' for end users and wants to know what to tell the executive sponsor. You have measured that time-to-first-token is 400 ms but full responses take 9 seconds because outputs are long. Which communication is most accurate and useful for the sponsor?
This separates the two latency components accurately: fast first token versus slow full completion, which is characteristic of long generations. It then offers two concrete, low-risk levers, streaming for perceived speed and output-length control for real speed, giving the sponsor a clear problem statement and a path forward rather than a vague promise.
Why this answer
Good stakeholder communication distinguishes between what users perceive and what the system actually measures. Fast time-to-first-token with slow completion is the signature of long outputs, so the honest message names that cause and pairs it with remedies that match it: streaming to improve perceived responsiveness and output-length guidance to reduce real duration. Blaming infrastructure, swapping models without evidence, or declaring the problem nonexistent all misrepresent the measurements.
Exam trap
The trap here is collapsing all latency into a single number and choosing a remedy, such as a model swap or infrastructure escalation, that does not correspond to the measured bottleneck.