Schema.ai· vLLM
vLLM official site Log in
All versions · Version range: v0.4.2

In vLLM v0.4.2 SchedulerConfig, when max_num_batched_tokens is unset and enable_chunked_prefill=True, code sets max_num_batched_tokens to 512.

SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code

Setting

max_num_batched_tokens

Applies when

max_num_batched_tokens is unsetenable_chunked_prefill=True

Evidence

vLLM v0.4.2 config.py(vllm/config.py:615-622)
definesprimarysource_codecode_introspection

Open cited lines in immutable upstream source ↗

615:         if max_num_batched_tokens is not None:
616:             self.max_num_batched_tokens = max_num_batched_tokens
617:         else:
618:             if enable_chunked_prefill:
619:                 # It is the values that have the best balance between ITL
620:                 # and TTFT on A100. Note it is not optimized for throughput.
621:                 self.max_num_batched_tokens = 512
622:             else:

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • defaults in later vLLM releases
  • defaults when enable_chunked_prefill is false

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.