All versions
vLLM latest Optimization and Tuning snapshot
- Type
- docs_page
- Publisher
- vllm-project
- Relationship
- first_party
- Version
- vLLM main-2026-07-21
- Retrieved
- Jul 21
- Local snapshot
- corpus/domains/vllm/core-v0/evidence/snapshots/vllm-docs-latest-optimization-20260721.md
- SHA-256
- 86ca9cde5064d129…
Claims citing this source (4)
| Claim | Locator | Verified |
|---|---|---|
The latest optimization docs snapshot says V1 enables chunked prefill by default whenever possible and uses max_num_batched_tokens as the token budget for pending prefills. | docs/configuration/optimization.md:53-53 | Aug 20 |
The latest optimization docs snapshot recommends max_num_batched_tokens greater than 8192 for optimal throughput, especially for smaller models on large GPUs. | docs/configuration/optimization.md:62-67 | Aug 20 |
The latest optimization docs snapshot warns that when enable_chunked_prefill=False, max_num_batched_tokens must be greater than max_model_len or vLLM may crash at startup. | docs/configuration/optimization.md:70-71 | Aug 20 |
The latest optimization docs snapshot says that decreasing max_num_seqs can reduce concurrent requests in a batch and require less KV cache space when preemptions are frequent. | docs/configuration/optimization.md:37-42 | Aug 20 |