Schema.ai· vLLM· Schema 0.2.1
vLLM official site vLLM v0.23.0Log in

#38194: Prefix-cache percentages require comparable definitions

Two hit rates are not comparable until their populations and scopes match

GitHub #38194retrievalhonestyvLLM v0.23.0

Reported environment

High-concurrency multi-turn workload comparing vLLM with a hosted API; no vLLM version reported.

The nuance

A percentage hides its numerator, denominator, labels, process scope, routing, reset history, tokenization, and time window.

Expected behavior

Give the vLLM token-rate formula and comparison preconditions; do not diagnose a regression from an undocumented percentage.

What Schema establishes

Schema defines the vLLM side; the other system and matched runtime data are required for comparison.

Runtime evidence still needed

  • Both formulas and scopes
  • Identical rendered token streams and routing
  • Matched warm-up and windows

Try the questions (2)

Each question runs live against the product.

Why is vLLM prefix cache hit rate lower than another providerrun →Can I compare vLLM and provider cache percentages directlyrun →

Grounding claims

Topics, entities, and gaps