seed slicecovered at v0.26.0legacy evidence: v0.23.0evidence: source code and documentation
Coverage
Coverage for the vLLM v0.23.0 first-party prefix-cache counters and how to compute and interpret the prefix cache hit rate.
Strong for what vllm:prefix_cache_queries and vllm:prefix_cache_hits count (tokens), where they are set from, and how to compute hit rate; none for healthy hit-rate thresholds or workload-specific tuning.
Scoped to the two first-party prefix-cache counters; external/KV-connector prefix-cache counters are inventoried but not semantically covered.
Coverage at vLLM v0.26.0
Every declared topic entity has applicable claims, availability, or admitted history at this version.
0 claims2 history facts2/2 entities
History-backed metrics: prefix_cache_hits, prefix_cache_queries
Claims in this topic (0)
No claims at this version scope.
Known gaps
The corpus has no benchmark or operator evidence for healthy prefix-cache hit-rate ranges; hit rate is workload-dependent (prompt-prefix sharing).
Schema.ai can explain how to compute the hit rate but must not grade an observed value as good or bad.
Check: Add benchmark or design-partner telemetry evidence before serving hit-rate thresholds.
vllm:external_prefix_cache_queries/hits (KV-connector prefix cache) are availability-inventoried but carry no semantic claims.
Existence answers are available for the external counters; interpretation claims are not.
Check: Promote KV-connector prefix-cache semantics from the connector metrics source when connector deployments become an eval target.