In vLLM v0.23.0, vllm:request_time_per_output_token_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.
MetricvLLM v0.23.0confidence: highverified 1 hr agoverified-code
Metric
request_time_per_output_token_secondsApplies when
PrometheusStatLogger metrics publisherV1 metrics
Evidence
vLLM v0.23.0 v1 metrics loggers.py(vllm/v1/metrics/loggers.py:817-842)
definesprimarysource_codecode_introspection
Does not establish
- whether these boundaries suit any particular workload
- the distribution of observed values
Guards
vllm_version= v0.23.0metric_name=vllm:request_time_per_output_token_seconds