In vLLM v0.23.0, vllm:cache_config_info is registered as a Prometheus gauge. | v0.23.0 | Metric | cache_config_info | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:corrupted_requests is registered as a Prometheus counter when `VLLM_COMPUTE_NANS_IN_LOGITS=1`. | v0.23.0 | Metric | corrupted_requests | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:corrupted_requests is registered with the label names model_name, engine. | v0.23.0 | Metric | corrupted_requests | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | e2e_request_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:e2e_request_latency_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | e2e_request_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:e2e_request_latency_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0. | v0.23.0 | Metric | e2e_request_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:engine_sleep_state is registered as a Prometheus gauge. | v0.23.0 | Metric | engine_sleep_state | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:engine_sleep_state is registered with the label names model_name, engine, sleep_state. | v0.23.0 | Metric | engine_sleep_state | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:estimated_flops_per_gpu_total is registered as a Prometheus counter. | v0.23.0 | Metric | estimated_flops_per_gpu_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:estimated_read_bytes_per_gpu_total is registered as a Prometheus counter. | v0.23.0 | Metric | estimated_read_bytes_per_gpu_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:estimated_write_bytes_per_gpu_total is registered as a Prometheus counter. | v0.23.0 | Metric | estimated_write_bytes_per_gpu_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered as a Prometheus counter. | v0.23.0 | Metric | external_prefix_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:external_prefix_cache_hits is registered with the label names model_name, engine. | v0.23.0 | Metric | external_prefix_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered as a Prometheus counter. | v0.23.0 | Metric | external_prefix_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:external_prefix_cache_queries is registered with the label names model_name, engine. | v0.23.0 | Metric | external_prefix_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:generation_tokens is registered as a Prometheus counter. | v0.23.0 | Metric | generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:generation_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | inter_token_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:inter_token_latency_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | inter_token_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:inter_token_latency_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0. | v0.23.0 | Metric | inter_token_latency_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:iteration_tokens_total is registered as a Prometheus histogram. | v0.23.0 | Metric | iteration_tokens_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:iteration_tokens_total is registered with the label names model_name, engine. | v0.23.0 | Metric | iteration_tokens_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:iteration_tokens_total declares 13 literal histogram bucket boundaries in its registration, from 1.0 to 16384.0. | v0.23.0 | Metric | iteration_tokens_total | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`. | v0.23.0 | Metric | kv_block_idle_before_evict_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | kv_block_idle_before_evict_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_idle_before_evict_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0. | v0.23.0 | Metric | kv_block_idle_before_evict_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`. | v0.23.0 | Metric | kv_block_lifetime_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | kv_block_lifetime_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_lifetime_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0. | v0.23.0 | Metric | kv_block_lifetime_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered as a Prometheus histogram when `observability_config.kv_cache_metrics=True`. | v0.23.0 | Metric | kv_block_reuse_gap_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | kv_block_reuse_gap_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_block_reuse_gap_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.001 to 1800.0. | v0.23.0 | Metric | kv_block_reuse_gap_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_cache_usage_perc is registered with the label names model_name, engine. | v0.23.0 | Metric | kv_cache_usage_perc | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:lora_requests_info is registered as a Prometheus gauge when `lora_config is not None`. | v0.23.0 | Metric | lora_requests_info | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:mm_cache_hits is registered as a Prometheus counter. | v0.23.0 | Metric | mm_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:mm_cache_hits is registered with the label names model_name, engine. | v0.23.0 | Metric | mm_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:mm_cache_queries is registered as a Prometheus counter. | v0.23.0 | Metric | mm_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:mm_cache_queries is registered with the label names model_name, engine. | v0.23.0 | Metric | mm_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_preemptions is registered with the label names model_name, engine. | v0.23.0 | Metric | num_preemptions | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_running is registered with the label names model_name, engine. | v0.23.0 | Metric | num_requests_running | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_waiting is registered with the label names model_name, engine. | v0.23.0 | Metric | num_requests_waiting | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is registered with the label names model_name, engine, reason. | v0.23.0 | Metric | num_requests_waiting_by_reason | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_hits is registered with the label names model_name, engine. | v0.23.0 | Metric | prefix_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_queries is registered with the label names model_name, engine. | v0.23.0 | Metric | prefix_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens is registered as a Prometheus counter. | v0.23.0 | Metric | prompt_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | prompt_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered as a Prometheus counter. | v0.23.0 | Metric | prompt_tokens_by_source | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens_by_source is registered with the label names model_name, engine, source. | v0.23.0 | Metric | prompt_tokens_by_source | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered as a Prometheus counter. | v0.23.0 | Metric | prompt_tokens_cached | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prompt_tokens_cached is registered with the label names model_name, engine. | v0.23.0 | Metric | prompt_tokens_cached | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | request_decode_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_decode_time_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | request_decode_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_decode_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0. | v0.23.0 | Metric | request_decode_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_generation_tokens is registered as a Prometheus histogram. | v0.23.0 | Metric | request_generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_generation_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | request_generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | request_inference_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_inference_time_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | request_inference_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_inference_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0. | v0.23.0 | Metric | request_inference_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered as a Prometheus histogram. | v0.23.0 | Metric | request_max_num_generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_max_num_generation_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | request_max_num_generation_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_params_max_tokens is registered as a Prometheus histogram. | v0.23.0 | Metric | request_params_max_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_params_max_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | request_params_max_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_params_n is registered as a Prometheus histogram. | v0.23.0 | Metric | request_params_n | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_params_n is registered with the label names model_name, engine. | v0.23.0 | Metric | request_params_n | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_params_n declares 5 literal histogram bucket boundaries in its registration, from 1.0 to 20.0. | v0.23.0 | Metric | request_params_n | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered as a Prometheus histogram. | v0.23.0 | Metric | request_prefill_kv_computed_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prefill_kv_computed_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | request_prefill_kv_computed_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | request_prefill_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prefill_time_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | request_prefill_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prefill_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0. | v0.23.0 | Metric | request_prefill_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prompt_tokens is registered as a Prometheus histogram. | v0.23.0 | Metric | request_prompt_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_prompt_tokens is registered with the label names model_name, engine. | v0.23.0 | Metric | request_prompt_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | request_queue_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_queue_time_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | request_queue_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_queue_time_seconds declares 21 literal histogram bucket boundaries in its registration, from 0.3 to 7680.0. | v0.23.0 | Metric | request_queue_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_success is registered as a Prometheus counter. | v0.23.0 | Metric | request_success | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_success is registered with the label names model_name, engine, finished_reason. | v0.23.0 | Metric | request_success | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | request_time_per_output_token_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | request_time_per_output_token_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:request_time_per_output_token_seconds declares 19 literal histogram bucket boundaries in its registration, from 0.01 to 80.0. | v0.23.0 | Metric | request_time_per_output_token_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens is registered as a Prometheus counter when `speculative_config is not None`. | v0.23.0 | Metric | spec_decode_num_accepted_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:spec_decode_num_accepted_tokens_per_pos is registered as a Prometheus counter when `speculative_config is not None`. | v0.23.0 | Metric | spec_decode_num_accepted_tokens_per_pos | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:spec_decode_num_draft_tokens is registered as a Prometheus counter when `speculative_config is not None`. | v0.23.0 | Metric | spec_decode_num_draft_tokens | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:spec_decode_num_drafts is registered as a Prometheus counter when `speculative_config is not None`. | v0.23.0 | Metric | spec_decode_num_drafts | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered as a Prometheus histogram. | v0.23.0 | Metric | time_to_first_token_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:time_to_first_token_seconds is registered with the label names model_name, engine. | v0.23.0 | Metric | time_to_first_token_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:time_to_first_token_seconds declares 22 literal histogram bucket boundaries in its registration, from 0.001 to 2560.0. | v0.23.0 | Metric | time_to_first_token_seconds | 7 hr ago | github.com |
For each EngineCoreEventType.QUEUED event processed by IterationStats.update_from_events, req_stats.queued_ts is set to event.timestamp. | v0.23.0 | Metric | request_queue_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, decode_time is computed as req_stats.last_token_ts - req_stats.first_token_ts; the source comment describes the decode interval as from first NEW_TOKEN to last NEW_TOKEN and says any preemptions during decode are included. | v0.23.0 | Metric | request_decode_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, queued_time is computed as req_stats.scheduled_ts - req_stats.queued_ts; the source comment describes the queued interval as from first QUEUED event to first SCHEDULED. | v0.23.0 | Metric | request_queue_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, prefill_time is computed as req_stats.first_token_ts - req_stats.scheduled_ts; the source comment describes the prefill interval as from first SCHEDULED to first NEW_TOKEN and says any preemptions during prefill are included in the interval. | v0.23.0 | Metric | request_prefill_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, inference_time is computed as req_stats.last_token_ts - req_stats.scheduled_ts; the source comment describes the inference interval as from first SCHEDULED to last NEW_TOKEN and says any preemptions during prefill or decode are included. | v0.23.0 | Metric | request_inference_time_seconds | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_queries is a Prometheus counter measured in queried tokens — it counts tokens, not requests or cache blocks. | v0.23.0 | Metric | prefix_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_hits is a Prometheus counter measured in cached (hit) tokens — it counts tokens, not requests or cache blocks. | v0.23.0 | Metric | prefix_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_queries is incremented from SchedulerStats.prefix_cache_stats.queries on each logging update and observes prefix-cache token reuse for the first-party prefix cache (KV-connector prefix-cache reuse is reported by the separate vllm:external_prefix_cache_* counters). | v0.23.0 | Metric | prefix_cache_queries | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:prefix_cache_hits is incremented from SchedulerStats.prefix_cache_stats.hits on each logging update and observes prefix-cache token reuse for the first-party prefix cache (KV-connector prefix-cache reuse is reported by the separate vllm:external_prefix_cache_* counters). | v0.23.0 | Metric | prefix_cache_hits | 7 hr ago | github.com |
The vLLM v0.23.0 metrics design docs state the metric of interest is the prefix cache hit rate (hits per query): the counters are exposed raw so operators compute the rate over an interval of their choosing with PromQL (rate(hits)/rate(queries)); vLLM's own logging equivalent aggregates hit_rate over the most recent queries. | v0.23.0 | Metric | prefix_cache_hits | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_running is a Prometheus gauge for the number of requests in model execution batches. | v0.23.0 | Metric | num_requests_running | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_running is set from SchedulerStats.num_running_reqs and observes scheduler running-request state. | v0.23.0 | Metric | num_requests_running | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_waiting is a Prometheus gauge for the number of requests waiting to be processed. | v0.23.0 | Metric | num_requests_waiting | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_requests_waiting_by_reason is a Prometheus gauge with reason labels capacity and deferred, and its reasons sum to vllm:num_requests_waiting. | v0.23.0 | Metric | num_requests_waiting_by_reason | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_cache_usage_perc is a Prometheus gauge for KV-cache usage where 1 means 100 percent usage. | v0.23.0 | Metric | kv_cache_usage_perc | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:kv_cache_usage_perc is set from SchedulerStats.kv_cache_usage and observes KV-cache usage pressure. | v0.23.0 | Metric | kv_cache_usage_perc | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_preemptions is a Prometheus counter for cumulative preemptions from the engine. | v0.23.0 | Metric | num_preemptions | 7 hr ago | github.com |
In vLLM v0.23.0, vllm:num_preemptions increments from IterationStats.num_preempted_reqs, which is increased when engine-core events mark requests as preempted. | v0.23.0 | Metric | num_preemptions | 7 hr ago | github.com |