All versions · Version range: v0.4.2
In vLLM v0.4.2, SchedulerConfig raises an error when max_num_batched_tokens is smaller than max_model_len and enable_chunked_prefill=False.
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code
Setting
max_num_batched_tokensApplies when
enable_chunked_prefill=False
Evidence
vLLM v0.4.2 config.py(vllm/config.py:639-647)
definesprimarysource_codecode_introspection
639: if (self.max_num_batched_tokens < self.max_model_len
640: and not self.chunked_prefill_enabled):
641: raise ValueError(
642: f"max_num_batched_tokens ({self.max_num_batched_tokens}) is "
643: f"smaller than max_model_len ({self.max_model_len}). "
644: "This effectively limits the maximum sequence length to "
645: "max_num_batched_tokens and makes vLLM reject longer "
646: "sequences. Please increase max_num_batched_tokens or "
647: "decrease max_model_len.")Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- the exact error behavior in later vLLM versions