All versions · Version range: v0.4.2
In vLLM v0.4.2, max_num_seqs is exposed through the SchedulerConfig.__init__ Python parameter max_num_seqs.
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code
Setting
max_num_seqsApplies when
SchedulerConfig construction
Evidence
vLLM v0.4.2 config.py(vllm/config.py:584-609)
definesprimarysource_codecode_introspection
584: class SchedulerConfig: 585: """Scheduler configuration. 586: 587: Args: 588: max_num_batched_tokens: Maximum number of tokens to be processed in 589: a single iteration. 590: max_num_seqs: Maximum number of sequences to be processed in a single 591: iteration. 592: max_model_len: Maximum length of a sequence (including prompt 593: and generated text). 594: use_v2_block_manager: Whether to use the BlockSpaceManagerV2 or not. 595: num_lookahead_slots: The number of slots to allocate per sequence per 596: step, beyond the known token ids. This is used in speculative 597: decoding to store KV activations of tokens which may or may not be 598: accepted. 599: delay_factor: Apply a delay (of delay factor multiplied by previous 600: prompt latency) before scheduling next prompt. 601: enable_chunked_prefill: If True, prefill requests can be chunked based 602: on the remaining max_num_batched_tokens. 603: """ 604: 605: def __init__( 606: self, 607: max_num_batched_tokens: Optional[int], 608: max_num_seqs: int, 609: max_model_len: int,
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- the CLI default for
max_num_seqs - defaults in later vLLM releases
- how engine argument parsing sets
max_num_seqsbeforeSchedulerConfigconstruction