All versions · Version range: v0.4.2
In vLLM v0.4.2 SchedulerConfig, when max_num_batched_tokens is unset and enable_chunked_prefill=True, code sets max_num_batched_tokens to 512.
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code
Setting
max_num_batched_tokensApplies when
max_num_batched_tokens is unsetenable_chunked_prefill=True
Evidence
vLLM v0.4.2 config.py(vllm/config.py:615-622)
definesprimarysource_codecode_introspection
615: if max_num_batched_tokens is not None: 616: self.max_num_batched_tokens = max_num_batched_tokens 617: else: 618: if enable_chunked_prefill: 619: # It is the values that have the best balance between ITL 620: # and TTFT on A100. Note it is not optimized for throughput. 621: self.max_num_batched_tokens = 512 622: else:
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- defaults in later vLLM releases
- defaults when
enable_chunked_prefillis false