Schema.ai· vLLM· Schema 0.2.1
vLLM official site vLLM v0.23.0Log in

#40696: Full-block eligibility creates prefix-cache cliffs

A reusable prefix shorter than one effective block can still produce zero hits

GitHub #40696retrievalhonestyvLLM v0.23.0

Reported environment

Qwen3.5 hybrid/Mamba on A100 and H20 with a reported 528-token effective block.

The nuance

Full-block matching can create sharp eligibility boundaries; zero hits need not mean the counter is broken.

Expected behavior

Explain eligibility, connect cached/computed tokens to hit/query counters, and distinguish proposals from released behavior.

What Schema establishes

Semantics explain the mechanism; model geometry and net impact require version-matched measurement.

Runtime evidence still needed

  • Effective block geometry
  • Boundary sweep of cached and computed tokens plus QPS and latency

Try the questions (2)

Each question runs live against the product.

Why are prefix hits zero below a prompt boundaryrun →Why can prefix caching hurt around hybrid block sizerun →

Grounding claims

Topics, entities, and gaps