All versions · Version range: v0.4.2
In vLLM v0.4.2, SchedulerConfig raises an error when max_num_batched_tokens is smaller than max_num_seqs.
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code
Setting
max_num_seqsApplies when
SchedulerConfig argument verification
Evidence
vLLM v0.4.2 config.py(vllm/config.py:649-653)
definesprimarysource_codecode_introspection
649: if self.max_num_batched_tokens < self.max_num_seqs:
650: raise ValueError(
651: f"max_num_batched_tokens ({self.max_num_batched_tokens}) must "
652: "be greater than or equal to max_num_seqs "
653: f"({self.max_num_seqs}).")Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- constraints in later vLLM releases
- that
max_num_seqsalone determines token budget