Schema.ai· vLLM
vLLM official site Log in
All versions

The latest optimization docs snapshot says V1 enables chunked prefill by default whenever possible and uses max_num_batched_tokens as the token budget for pending prefills.

BehaviorvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs

Entity

chunked prefill decode-priority scheduling

Applies when

V1chunked prefill available

Evidence

vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:53-53)
documentsprimarydocumentationdocs_extraction

Open cited lines in immutable upstream source ↗

53: In V1, **chunked prefill is enabled by default whenever possible**. With chunked prefill enabled, the scheduling policy prioritizes decode requests. It batches all pending decode requests before scheduling any prefill operations. When there are available tokens in the `max_num_batched_tokens` budget, it schedules pending prefills. If a pending prefill request cannot fit into `max_num_batched_tokens`, it automatically chunks it.

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • which released vLLM tag first introduced this behavior
  • runtime defaults for every deployment

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.