All versions · Version range: v0.4.2
In vLLM v0.4.2 SchedulerConfig, when max_num_batched_tokens is unset and enable_chunked_prefill=False, code sets max_num_batched_tokens to max(max_model_len, 2048).
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code
Setting
max_num_batched_tokensApplies when
max_num_batched_tokens is unsetenable_chunked_prefill=False
Evidence
vLLM v0.4.2 config.py(vllm/config.py:623-626)
definesprimarysource_codecode_introspection
623: # If max_model_len is too short, use 2048 as the default value 624: # for higher throughput. 625: self.max_num_batched_tokens = max(max_model_len, 2048) 626: if enable_chunked_prefill:
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- defaults in later vLLM releases
- the value for deployments that explicitly set
max_num_batched_tokens