Schema.ai· vLLM
vLLM official site Log in
All versions · Version range: v0.4.2

In vLLM v0.4.2, max_num_seqs is exposed through the SchedulerConfig.__init__ Python parameter max_num_seqs.

SettingvLLM v0.4.2confidence: highverified 1 day agoverified-code

Setting

max_num_seqs

Applies when

SchedulerConfig construction

Evidence

vLLM v0.4.2 config.py(vllm/config.py:584-609)
definesprimarysource_codecode_introspection

Open cited lines in immutable upstream source ↗

584: class SchedulerConfig:
585:     """Scheduler configuration.
586: 
587:     Args:
588:         max_num_batched_tokens: Maximum number of tokens to be processed in
589:             a single iteration.
590:         max_num_seqs: Maximum number of sequences to be processed in a single
591:             iteration.
592:         max_model_len: Maximum length of a sequence (including prompt
593:             and generated text).
594:         use_v2_block_manager: Whether to use the BlockSpaceManagerV2 or not.
595:         num_lookahead_slots: The number of slots to allocate per sequence per
596:             step, beyond the known token ids. This is used in speculative
597:             decoding to store KV activations of tokens which may or may not be
598:             accepted.
599:         delay_factor: Apply a delay (of delay factor multiplied by previous
600:             prompt latency) before scheduling next prompt.
601:         enable_chunked_prefill: If True, prefill requests can be chunked based
602:             on the remaining max_num_batched_tokens.
603:     """
604: 
605:     def __init__(
606:         self,
607:         max_num_batched_tokens: Optional[int],
608:         max_num_seqs: int,
609:         max_model_len: int,

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • the CLI default for max_num_seqs
  • defaults in later vLLM releases
  • how engine argument parsing sets max_num_seqs before SchedulerConfig construction

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.