All versions
The latest optimization docs snapshot says V1 enables chunked prefill by default whenever possible and uses max_num_batched_tokens as the token budget for pending prefills.
BehaviorvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs
Entity
chunked prefill decode-priority schedulingApplies when
V1chunked prefill available
Evidence
vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:53-53)
documentsprimarydocumentationdocs_extraction
53: In V1, **chunked prefill is enabled by default whenever possible**. With chunked prefill enabled, the scheduling policy prioritizes decode requests. It batches all pending decode requests before scheduling any prefill operations. When there are available tokens in the `max_num_batched_tokens` budget, it schedules pending prefills. If a pending prefill request cannot fit into `max_num_batched_tokens`, it automatically chunks it.
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- which released vLLM tag first introduced this behavior
- runtime defaults for every deployment