All versions
The latest optimization docs snapshot warns that when enable_chunked_prefill=False, max_num_batched_tokens must be greater than max_model_len or vLLM may crash at startup.
SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs
Setting
max_num_batched_tokensApplies when
enable_chunked_prefill=False
Evidence
vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:70-71)
documentsprimarydocumentationdocs_extraction
70: When chunked prefill is disabled, `max_num_batched_tokens` must be greater than `max_model_len`. 71: In that case, if `max_num_batched_tokens < max_model_len`, vLLM may crash at server start‑up.
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- the exact exception path in a released version