Schema.ai· vLLM· Schema 0.2.0
vLLM official site vLLM v0.23.0Log in

Claims

ClaimVersionAboutEntityVerifiedSources
In vLLM v0.23.0, vllm:cache_config_info is registered as a Prometheus gauge.v0.23.0Metriccache_config_info7 hr agogithub.com
In vLLM v0.23.0, vllm:corrupted_requests is registered as a Prometheus counter when `VLLM_COMPUTE_NANS_IN_LOGITS=1`.v0.23.0Metriccorrupted_requests7 hr agogithub.com
In vLLM v0.23.0, vllm:corrupted_requests is registered with the label names model_name, engine.v0.23.0Metriccorrupted_requests7 hr agogithub.com
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered as a Prometheus histogram.v0.23.0Metrice2e_request_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered with the label names model_name, engine.v0.23.0Metrice2e_request_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:e2e_request_latency_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.v0.23.0Metrice2e_request_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:engine_sleep_state is registered as a Prometheus gauge.v0.23.0Metricengine_sleep_state7 hr agogithub.com
In vLLM v0.23.0, vllm:engine_sleep_state is registered with the label names model_name, engine, sleep_state.v0.23.0Metricengine_sleep_state7 hr agogithub.com
In vLLM v0.23.0, vllm:estimated_flops_per_gpu_total is registered as a Prometheus counter.v0.23.0Metricestimated_flops_per_gpu_total7 hr agogithub.com
In vLLM v0.23.0, vllm:estimated_read_bytes_per_gpu_total is registered as a Prometheus counter.v0.23.0Metricestimated_read_bytes_per_gpu_total7 hr agogithub.com
In vLLM v0.23.0, vllm:estimated_write_bytes_per_gpu_total is registered as a Prometheus counter.v0.23.0Metricestimated_write_bytes_per_gpu_total7 hr agogithub.com
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered as a Prometheus counter.v0.23.0Metricexternal_prefix_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered with the label names model_name, engine.v0.23.0Metricexternal_prefix_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered as a Prometheus counter.v0.23.0Metricexternal_prefix_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered with the label names model_name, engine.v0.23.0Metricexternal_prefix_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:generation_tokens is registered as a Prometheus counter.v0.23.0Metricgeneration_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:generation_tokens is registered with the label names model_name, engine.v0.23.0Metricgeneration_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered as a Prometheus histogram.v0.23.0Metricinter_token_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered with the label names model_name, engine.v0.23.0Metricinter_token_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:inter_token_latency_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.v0.23.0Metricinter_token_latency_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:iteration_tokens_total is registered as a Prometheus histogram.v0.23.0Metriciteration_tokens_total7 hr agogithub.com
In vLLM v0.23.0, vllm:iteration_tokens_total is registered with the label names model_name, engine.v0.23.0Metriciteration_tokens_total7 hr agogithub.com
In vLLM v0.23.0, vllm:iteration_tokens_total declares 13 literal histogram bucket boundaries in its registration, from 1.0 to 16384.0.v0.23.0Metriciteration_tokens_total7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`.v0.23.0Metrickv_block_idle_before_evict_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered with the label names model_name, engine.v0.23.0Metrickv_block_idle_before_evict_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.v0.23.0Metrickv_block_idle_before_evict_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`.v0.23.0Metrickv_block_lifetime_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered with the label names model_name, engine.v0.23.0Metrickv_block_lifetime_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.v0.23.0Metrickv_block_lifetime_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`.v0.23.0Metrickv_block_reuse_gap_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered with the label names model_name, engine.v0.23.0Metrickv_block_reuse_gap_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0.v0.23.0Metrickv_block_reuse_gap_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_cache_usage_perc is registered with the label names model_name, engine.v0.23.0Metrickv_cache_usage_perc7 hr agogithub.com
In vLLM v0.23.0, vllm:lora_requests_info is registered as a Prometheus gauge when `lora_config is not None`.v0.23.0Metriclora_requests_info7 hr agogithub.com
In vLLM v0.23.0, vllm:mm_cache_hits is registered as a Prometheus counter.v0.23.0Metricmm_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:mm_cache_hits is registered with the label names model_name, engine.v0.23.0Metricmm_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:mm_cache_queries is registered as a Prometheus counter.v0.23.0Metricmm_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:mm_cache_queries is registered with the label names model_name, engine.v0.23.0Metricmm_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:num_preemptions is registered with the label names model_name, engine.v0.23.0Metricnum_preemptions7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_running is registered with the label names model_name, engine.v0.23.0Metricnum_requests_running7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_waiting is registered with the label names model_name, engine.v0.23.0Metricnum_requests_waiting7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is registered with the label names model_name, engine, reason.v0.23.0Metricnum_requests_waiting_by_reason7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_hits is registered with the label names model_name, engine.v0.23.0Metricprefix_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_queries is registered with the label names model_name, engine.v0.23.0Metricprefix_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens is registered as a Prometheus counter.v0.23.0Metricprompt_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens is registered with the label names model_name, engine.v0.23.0Metricprompt_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered as a Prometheus counter.v0.23.0Metricprompt_tokens_by_source7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered with the label names model_name, engine, source.v0.23.0Metricprompt_tokens_by_source7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered as a Prometheus counter.v0.23.0Metricprompt_tokens_cached7 hr agogithub.com
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered with the label names model_name, engine.v0.23.0Metricprompt_tokens_cached7 hr agogithub.com
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered as a Prometheus histogram.v0.23.0Metricrequest_decode_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered with the label names model_name, engine.v0.23.0Metricrequest_decode_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_decode_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.v0.23.0Metricrequest_decode_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_generation_tokens is registered as a Prometheus histogram.v0.23.0Metricrequest_generation_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_generation_tokens is registered with the label names model_name, engine.v0.23.0Metricrequest_generation_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered as a Prometheus histogram.v0.23.0Metricrequest_inference_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered with the label names model_name, engine.v0.23.0Metricrequest_inference_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_inference_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.v0.23.0Metricrequest_inference_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered as a Prometheus histogram.v0.23.0Metricrequest_max_num_generation_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered with the label names model_name, engine.v0.23.0Metricrequest_max_num_generation_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_params_max_tokens is registered as a Prometheus histogram.v0.23.0Metricrequest_params_max_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_params_max_tokens is registered with the label names model_name, engine.v0.23.0Metricrequest_params_max_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_params_n is registered as a Prometheus histogram.v0.23.0Metricrequest_params_n7 hr agogithub.com
In vLLM v0.23.0, vllm:request_params_n is registered with the label names model_name, engine.v0.23.0Metricrequest_params_n7 hr agogithub.com
In vLLM v0.23.0, vllm:request_params_n declares 5 literal histogram bucket boundaries in its registration, from 1.0 to 20.0.v0.23.0Metricrequest_params_n7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered as a Prometheus histogram.v0.23.0Metricrequest_prefill_kv_computed_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered with the label names model_name, engine.v0.23.0Metricrequest_prefill_kv_computed_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered as a Prometheus histogram.v0.23.0Metricrequest_prefill_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered with the label names model_name, engine.v0.23.0Metricrequest_prefill_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prefill_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.v0.23.0Metricrequest_prefill_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prompt_tokens is registered as a Prometheus histogram.v0.23.0Metricrequest_prompt_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_prompt_tokens is registered with the label names model_name, engine.v0.23.0Metricrequest_prompt_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered as a Prometheus histogram.v0.23.0Metricrequest_queue_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered with the label names model_name, engine.v0.23.0Metricrequest_queue_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_queue_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0.v0.23.0Metricrequest_queue_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_success is registered as a Prometheus counter.v0.23.0Metricrequest_success7 hr agogithub.com
In vLLM v0.23.0, vllm:request_success is registered with the label names model_name, engine, finished_reason.v0.23.0Metricrequest_success7 hr agogithub.com
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered as a Prometheus histogram.v0.23.0Metricrequest_time_per_output_token_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered with the label names model_name, engine.v0.23.0Metricrequest_time_per_output_token_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0.v0.23.0Metricrequest_time_per_output_token_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens is registered as a Prometheus counter when `speculative_config is not None`.v0.23.0Metricspec_decode_num_accepted_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens_per_pos is registered as a Prometheus counter when `speculative_config is not None`.v0.23.0Metricspec_decode_num_accepted_tokens_per_pos7 hr agogithub.com
In vLLM v0.23.0, vllm:spec_decode_num_draft_tokens is registered as a Prometheus counter when `speculative_config is not None`.v0.23.0Metricspec_decode_num_draft_tokens7 hr agogithub.com
In vLLM v0.23.0, vllm:spec_decode_num_drafts is registered as a Prometheus counter when `speculative_config is not None`.v0.23.0Metricspec_decode_num_drafts7 hr agogithub.com
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered as a Prometheus histogram.v0.23.0Metrictime_to_first_token_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered with the label names model_name, engine.v0.23.0Metrictime_to_first_token_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:time_to_first_token_seconds declares 22 literal histogram bucket boundaries in its registration, from 0.001 to 2560.0.v0.23.0Metrictime_to_first_token_seconds7 hr agogithub.com
For each EngineCoreEventType.QUEUED event processed by IterationStats.update_from_events, req_stats.queued_ts is set to event.timestamp.v0.23.0Metricrequest_queue_time_seconds7 hr agogithub.com
In vLLM v0.23.0, decode_time is computed as req_stats.last_token_ts - req_stats.first_token_ts; the source comment describes the decode interval as from first NEW_TOKEN to last NEW_TOKEN and says any preemptions during decode are included.v0.23.0Metricrequest_decode_time_seconds7 hr agogithub.com
In vLLM v0.23.0, queued_time is computed as req_stats.scheduled_ts - req_stats.queued_ts; the source comment describes the queued interval as from first QUEUED event to first SCHEDULED.v0.23.0Metricrequest_queue_time_seconds7 hr agogithub.com
In vLLM v0.23.0, prefill_time is computed as req_stats.first_token_ts - req_stats.scheduled_ts; the source comment describes the prefill interval as from first SCHEDULED to first NEW_TOKEN and says any preemptions during prefill are included in the interval.v0.23.0Metricrequest_prefill_time_seconds7 hr agogithub.com
In vLLM v0.23.0, inference_time is computed as req_stats.last_token_ts - req_stats.scheduled_ts; the source comment describes the inference interval as from first SCHEDULED to last NEW_TOKEN and says any preemptions during prefill or decode are included.v0.23.0Metricrequest_inference_time_seconds7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_queries is a Prometheus counter measured in queried tokens it counts tokens, not requests or cache blocks.v0.23.0Metricprefix_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_hits is a Prometheus counter measured in cached (hit) tokens it counts tokens, not requests or cache blocks.v0.23.0Metricprefix_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_queries is incremented from SchedulerStats.prefix_cache_stats.queries on each logging update and observes prefix-cache token reuse for the first-party prefix cache (KV-connector prefix-cache reuse is reported by the separate vllm:external_prefix_cache_* counters).v0.23.0Metricprefix_cache_queries7 hr agogithub.com
In vLLM v0.23.0, vllm:prefix_cache_hits is incremented from SchedulerStats.prefix_cache_stats.hits on each logging update and observes prefix-cache token reuse for the first-party prefix cache (KV-connector prefix-cache reuse is reported by the separate vllm:external_prefix_cache_* counters).v0.23.0Metricprefix_cache_hits7 hr agogithub.com
The vLLM v0.23.0 metrics design docs state the metric of interest is the prefix cache hit rate (hits per query): the counters are exposed raw so operators compute the rate over an interval of their choosing with PromQL (rate(hits)/rate(queries)); vLLM's own logging equivalent aggregates hit_rate over the most recent queries.v0.23.0Metricprefix_cache_hits7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_running is a Prometheus gauge for the number of requests in model execution batches.v0.23.0Metricnum_requests_running7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_running is set from SchedulerStats.num_running_reqs and observes scheduler running-request state.v0.23.0Metricnum_requests_running7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_waiting is a Prometheus gauge for the number of requests waiting to be processed.v0.23.0Metricnum_requests_waiting7 hr agogithub.com
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is a Prometheus gauge with reason labels capacity and deferred, and its reasons sum to vllm:num_requests_waiting.v0.23.0Metricnum_requests_waiting_by_reason7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_cache_usage_perc is a Prometheus gauge for KV-cache usage where 1 means 100 percent usage.v0.23.0Metrickv_cache_usage_perc7 hr agogithub.com
In vLLM v0.23.0, vllm:kv_cache_usage_perc is set from SchedulerStats.kv_cache_usage and observes KV-cache usage pressure.v0.23.0Metrickv_cache_usage_perc7 hr agogithub.com
In vLLM v0.23.0, vllm:num_preemptions is a Prometheus counter for cumulative preemptions from the engine.v0.23.0Metricnum_preemptions7 hr agogithub.com
In vLLM v0.23.0, vllm:num_preemptions increments from IterationStats.num_preempted_reqs, which is increased when engine-core events mark requests as preempted.v0.23.0Metricnum_preemptions7 hr agogithub.com