All versions
The latest optimization docs snapshot says that decreasing max_num_seqs can reduce concurrent requests in a batch and require less KV cache space when preemptions are frequent.
SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs
Setting
max_num_seqsApplies when
frequent preemptionsKV cache pressure
Evidence
vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:37-42)
documentsprimarydocumentationdocs_extraction
37: While this mechanism ensures system robustness, preemption and recomputation can adversely affect end-to-end latency. 38: If you frequently encounter preemptions, consider the following actions: 39: 40: - Increase `gpu_memory_utilization`. vLLM pre-allocates GPU cache using this percentage of memory. By increasing utilization, you can provide more KV cache space. 41: - Decrease `max_num_seqs` or `max_num_batched_tokens`. This reduces the number of concurrent requests in a batch, thereby requiring less KV cache space. 42: - Increase `tensor_parallel_size`. This shards model weights across GPUs, allowing each GPU to have more memory available for KV cache. However, increasing this value may cause excessive synchronization overhead.
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- a universal recommended
max_num_seqsvalue - released-version applicability without checking the target tag
- that decreasing
max_num_seqsis always preferable to increasing GPU memory utilization or parallelism