Schema.ai· vLLM
vLLM official site Log in
All versions

The latest optimization docs snapshot says that decreasing max_num_seqs can reduce concurrent requests in a batch and require less KV cache space when preemptions are frequent.

SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs

Setting

max_num_seqs

Applies when

frequent preemptionsKV cache pressure

Evidence

vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:37-42)
documentsprimarydocumentationdocs_extraction

Open cited lines in immutable upstream source ↗

37: While this mechanism ensures system robustness, preemption and recomputation can adversely affect end-to-end latency.
38: If you frequently encounter preemptions, consider the following actions:
39: 
40: - Increase `gpu_memory_utilization`. vLLM pre-allocates GPU cache using this percentage of memory. By increasing utilization, you can provide more KV cache space.
41: - Decrease `max_num_seqs` or `max_num_batched_tokens`. This reduces the number of concurrent requests in a batch, thereby requiring less KV cache space.
42: - Increase `tensor_parallel_size`. This shards model weights across GPUs, allowing each GPU to have more memory available for KV cache. However, increasing this value may cause excessive synchronization overhead.

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • a universal recommended max_num_seqs value
  • released-version applicability without checking the target tag
  • that decreasing max_num_seqs is always preferable to increasing GPU memory utilization or parallelism

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.