In vLLM v0.23.0, vllm:request_time_per_output_token_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.
MetricvLLM v0.23.0confidence: highverified 14 min agoverified-code
Metric
request_time_per_output_token_secondsApplies when
PrometheusStatLogger metrics publisherV1 metrics
Evidence
vLLM v0.23.0 v1 metrics loggers.py(vllm/v1/metrics/loggers.py:817-842)
definesprimarysource_codecode_introspection
Does not establish
- whether these boundaries suit any particular workload
- the distribution of observed values
Guards
- vllm_version = v0.23.0
- metric_name = vllm:request_time_per_output_token_seconds