Schema.ai· vLLM
vLLM official site vLLM v0.23.0Log in

vLLM v0.23.0 metric identity

broadvLLM v0.23.0evidence: source code

Coverage

Identity facts for the complete v0.23.0 Prometheus metric inventory which metrics exist, their instrument kind, their label names, and their literal histogram bucket boundaries generated from the pinned metric registration code.

Strong for whether a fixed-name metric exists at v0.23.0 and what Prometheus instrument kind it is registered as, across the complete 45-metric inventory; strong for label names and literal bucket boundaries where registration declares them statically; none for units, for bucket boundaries computed at runtime, or for label sets built at runtime.

Generated slice, not hand-authored: every claim here is produced by scripts/generate_metric_identity_claims.py from registration call sites in the pinned v0.23.0 snapshots, and is reproducible by re-running it. Instrument-kind coverage is complete over the closed inventory (45 of 45); label and bucket coverage is partial by construction, and the shortfalls are declared below rather than inferred. Seven instrument-kind claims predate the generator, are hand-authored, and are checked against the extraction rather than overwritten by it.

Claims in this topic (94)

ClaimVerified
In vLLM v0.23.0, vllm:cache_config_info is registered as a Prometheus gauge.1 hr ago
In vLLM v0.23.0, vllm:corrupted_requests is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:corrupted_requests is registered as a Prometheus counter when `envs.VLLM_COMPUTE_NANS_IN_LOGITS` is enabled.1 hr ago
In vLLM v0.23.0, vllm:e2e_request_latency_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.1 hr ago
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:engine_sleep_state is registered with the label names model_name, engine, sleep_state.1 hr ago
In vLLM v0.23.0, vllm:engine_sleep_state is registered as a Prometheus gauge.1 hr ago
In vLLM v0.23.0, vllm:estimated_flops_per_gpu_total is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:estimated_read_bytes_per_gpu_total is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:estimated_write_bytes_per_gpu_total is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:generation_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:generation_tokens is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:inter_token_latency_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.1 hr ago
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:iteration_tokens_total declares 13 literal histogram bucket boundaries in its registration, from 1.0 to 16384.0.1 hr ago
In vLLM v0.23.0, vllm:iteration_tokens_total is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:iteration_tokens_total is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.1 hr ago
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered as a Prometheus histogram when `self.kv_cache_metrics_enabled` is enabled.1 hr ago
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.1 hr ago
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered as a Prometheus histogram when `self.kv_cache_metrics_enabled` is enabled.1 hr ago
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.1 hr ago
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered as a Prometheus histogram when `self.kv_cache_metrics_enabled` is enabled.1 hr ago
In vLLM v0.23.0, vllm:kv_cache_usage_perc is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:kv_cache_usage_perc is a Prometheus gauge for KV-cache usage where 1 means 100 percent usage.1 hr ago
In vLLM v0.23.0, vllm:lora_requests_info is registered as a Prometheus gauge when `vllm_config.lora_config is not None` is enabled.1 hr ago
In vLLM v0.23.0, vllm:mm_cache_hits is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:mm_cache_hits is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:mm_cache_queries is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:mm_cache_queries is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:num_preemptions is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:num_preemptions is a Prometheus counter for cumulative preemptions from the engine.1 hr ago
In vLLM v0.23.0, vllm:num_requests_running is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:num_requests_running is a Prometheus gauge for the number of requests in model execution batches.1 hr ago
In vLLM v0.23.0, vllm:num_requests_waiting is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:num_requests_waiting is a Prometheus gauge for the number of requests waiting to be processed.1 hr ago
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is registered with the label names model_name, engine, reason.1 hr ago
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is a Prometheus gauge with reason labels capacity and deferred, and its reasons sum to vllm:num_requests_waiting.1 hr ago
In vLLM v0.23.0, vllm:prefix_cache_hits is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:prefix_cache_hits is a Prometheus counter measured in cached (hit) tokens it counts tokens, not requests or cache blocks.1 hr ago
In vLLM v0.23.0, vllm:prefix_cache_queries is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:prefix_cache_queries is a Prometheus counter measured in queried tokens it counts tokens, not requests or cache blocks.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered with the label names model_name, engine, source.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:request_decode_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.1 hr ago
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_generation_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_generation_tokens is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_inference_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.1 hr ago
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_params_max_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_params_max_tokens is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_params_n declares 5 literal histogram bucket boundaries in its registration, from 1.0 to 20.0.1 hr ago
In vLLM v0.23.0, vllm:request_params_n is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_params_n is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_prefill_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.1 hr ago
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_prompt_tokens is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_prompt_tokens is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_queue_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.1 hr ago
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:request_success is registered with the label names model_name, engine, finished_reason.1 hr ago
In vLLM v0.23.0, vllm:request_success is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.1 hr ago
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered as a Prometheus histogram.1 hr ago
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens_per_pos is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:spec_decode_num_draft_tokens is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:spec_decode_num_drafts is registered as a Prometheus counter.1 hr ago
In vLLM v0.23.0, vllm:time_to_first_token_seconds declares 22 literal histogram bucket boundaries in its registration, from 0.001 to 2560.0.1 hr ago
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered with the label names model_name, engine.1 hr ago
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered as a Prometheus histogram.1 hr ago

Metrics in this topic (45)

Surfaces

Known gaps

The corpus records no unit of measure for any v0.23.0 metric. vLLM's registration code has no structured unit field `unit=` appears zero times in the pinned snapshots so the only in-code signal is prose inside the `documentation=` argument, which states a unit for a minority of metrics and would require inference for the rest.

Schema.ai can state a metric's Prometheus instrument kind but cannot answer "what unit is this measured in" from structured corpus data. Where a registration docstring states a unit, it may appear in a claim statement as source-faithful prose, but it is not a queryable field.

Check: Admit a has_unit predicate only together with a rule for what grounds a unit registration prose, the metrics design doc, or neither. Inferring from name suffixes such as `_seconds` is explicitly not grounded enough.

Five histogram metrics declare bucket boundaries computed at runtime via `build_1_2_5_buckets(max_model_len)` rather than as literal values; the corpus records no boundaries for them. Thirteen of the eighteen bucket declarations are literal and are recorded.

Schema.ai can list bucket boundaries for the thirteen histograms that declare them literally and must decline for the five computed ones. Asserting fixed boundaries for a set derived from a deployment's max_model_len would be false.

Check: Admit computed_from, or an equivalent way to state a derivation and its inputs, before recording any bucket fact for these five metrics.

Nine metrics build their label set at runtime rather than declaring it statically from a constructor parameter, from `metrics_info.keys()`, or from instance attributes so the corpus records no label names for them. Thirty-six of the forty-five have statically resolvable label sets.

Schema.ai can list label names for thirty-six metrics and must decline for the other nine, rather than guessing a label set from the common `model_name`/`engine` pattern.

Check: Resolving these requires reading construction sites outside the registration call. Decide whether cross-call-site resolution is in scope before recording labels for them.