Schema.ai· vLLM
vLLM official site Log in
All versions

The latest optimization docs snapshot warns that when enable_chunked_prefill=False, max_num_batched_tokens must be greater than max_model_len or vLLM may crash at startup.

SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs

Setting

max_num_batched_tokens

Applies when

enable_chunked_prefill=False

Evidence

vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:70-71)
documentsprimarydocumentationdocs_extraction

Open cited lines in immutable upstream source ↗

70:     When chunked prefill is disabled, `max_num_batched_tokens` must be greater than `max_model_len`.  
71:     In that case, if `max_num_batched_tokens < max_model_len`, vLLM may crash at server start‑up.

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • the exact exception path in a released version

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.