How should a latency budget be set?
Start with the user’s task and define the point where waiting becomes disruptive. A voice conversation has a tighter target than a report generated in the background. Then divide that end-to-end target among the stages that consume it. Every component inherits a limit from the experience it serves. For a voice turn, endpointing, transport, model work, tool calls, validation, and playback all count.
The numbers in the diagram are an example, not a universal prescription. The useful act is making the total and each allowance explicit. If model inference receives 300 milliseconds, the team can test whether the chosen model and serving setup meet that limit under real load.
Which latency measure matters?
Measure the event the user notices. Time to first token describes when text generation starts. A voice application may care about time to first audible response. A tool workflow may care about the confirmed result. The metric must cover the point where the interaction becomes useful, not the easiest timestamp to collect.
Report percentiles alongside the median. Averages hide the slow calls that damage trust. The p95 value shows a threshold that most requests meet while preserving visibility into the slower tail. Keep timeout and failure counts beside latency, because dropping slow requests can make the remaining sample look artificially fast.
How should teams measure it?
Add timestamps at every boundary through LLM observability. Record when input arrives, when retrieval or a tool begins and ends, when the model produces its first useful output, and when the client renders or plays it. A component can meet its local target while the full interaction still misses the budget. Shared queues, retries, serialization, and client buffering often explain the missing time.
Test streaming inference, tool calls, and recovery under representative concurrency and network conditions. A realtime AI API should expose event timestamps so delay can be assigned to the correct stage instead of guessed from one total duration.
What happens when the budget is missed?
Decide the fallback before production. The system may choose a faster model, skip a nonessential enrichment step, return a partial result, continue in the background, or hand the task to a person. A latency fallback must preserve correctness and permissions. Speed is not permission to omit a required check or repeat a completed action.
Optimize the stage that actually consumes the tail, not the stage that is easiest to rewrite. Re-measure after every change. Moving 50 milliseconds from model work into a slower network hop has not improved the user experience, even if one service dashboard turns green.