6 matching claims
Each result is a source-grounded claim — click through for its evidence, guards, and graph context. Ranked by: Lexical overlap against claim text, object values, applicability, and entity names.
1Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230source.vllm.code.v1_metrics_stats.v0230source.vllm.code.v1_metrics_loggers.v0230
Does not establish (2)
- why reuse is high or low for a given workload
- connector/external prefix-cache reuse
Matched query terms against entity.vllm.metric.prefix_cache_queries and v0.23.0.
2Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230source.vllm.code.v1_metrics_stats.v0230source.vllm.code.v1_metrics_loggers.v0230
Does not establish (2)
- why reuse is high or low for a given workload
- connector/external prefix-cache reuse
Matched query terms against entity.vllm.metric.prefix_cache_hits and v0.23.0.
3Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.docs.design_metrics.v0230
Does not establish (2)
- a healthy or target hit rate (workload-dependent)
- how to increase hit rate for a given workload
Matched query terms against entity.vllm.metric.prefix_cache_hits and v0.23.0.
4Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230source.vllm.code.v1_metrics_loggers.v0230
Does not establish (3)
- a healthy hit-rate threshold
- request-level cache behavior (the unit is tokens, not requests)
- cache sizing guidance
Matched query terms against entity.vllm.metric.prefix_cache_queries and v0.23.0.
5Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230source.vllm.code.v1_metrics_loggers.v0230
Does not establish (3)
- a healthy hit-rate threshold
- request-level cache behavior (the unit is tokens, not requests)
- cache sizing guidance
Matched query terms against entity.vllm.metric.prefix_cache_hits and v0.23.0.
6Claimcaveatconfidence: highfreshness: fresh · full
Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230
Does not establish (3)
- how much free KV cache remains for a specific model
- whether a workload will preempt
- a universal alert threshold
Matched query terms against entity.vllm.metric.kv_cache_usage_perc and v0.23.0.