All versions
The latest optimization docs snapshot recommends max_num_batched_tokens greater than 8192 for optimal throughput, especially for smaller models on large GPUs.
SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs
Setting
max_num_batched_tokensApplies when
smaller modelslarge GPUsthroughput optimization
Evidence
vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:62-67)
documentsprimarydocumentationdocs_extraction
62: You can tune the performance by adjusting `max_num_batched_tokens`: 63: 64: - Smaller values (e.g., 2048) achieve better ITL because there are fewer prefills slowing down decodes. 65: - Higher values achieve better time to first token (TTFT) as you can process more prefill tokens in a batch. 66: - For optimal throughput, we recommend setting `max_num_batched_tokens > 8192` especially for smaller models on large GPUs. 67: - If `max_num_batched_tokens` is the same as `max_model_len`, that's almost the equivalent to the V0 default scheduling policy (except that it still prioritizes decodes).
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- a universal default value
- latency-optimal settings
- released-version applicability without checking the target tag