Schema.ai· vLLM· Schema 0.2.1
vLLM official site vLLM v0.23.0Log in

#35048: Aggregate throughput can hide experience regressions

Equal tokens per second does not imply equal TTFT, queueing, or decode pace

GitHub #35048retrievalhonestyvLLM v0.23.0

Reported environment

Same model and hardware across an upgrade with similar throughput but worse TTFT and output rate.

The nuance

Token and iteration metrics describe system work; TTFT queue ITL TPOT and E2E describe distinct parts of per-request experience.

Expected behavior

Return both families and require a fixed workload/version-matched comparison before assigning cause.

What Schema establishes

Schema supplies the decomposition; root cause requires controlled A/B or commit bisect.

Runtime evidence still needed

  • Fixed matched workload environment and dependencies
  • Same-window throughput and per-request latency distributions

Try the questions (2)

Each question runs live against the product.

How can throughput stay constant while inter-token latency worsensrun →Which metrics separate aggregate efficiency from user experiencerun →

Grounding claims

Topics, entities, and gaps