Schema.ai· vLLM
vLLM official site Log in
All versions

The latest optimization docs snapshot recommends max_num_batched_tokens greater than 8192 for optimal throughput, especially for smaller models on large GPUs.

SettingvLLM main-2026-07-21confidence: mediumverified Aug 20verified-docs

Setting

max_num_batched_tokens

Applies when

smaller modelslarge GPUsthroughput optimization

Evidence

vLLM latest Optimization and Tuning snapshot(docs/configuration/optimization.md:62-67)
documentsprimarydocumentationdocs_extraction

Open cited lines in immutable upstream source ↗

62: You can tune the performance by adjusting `max_num_batched_tokens`:
63: 
64: - Smaller values (e.g., 2048) achieve better ITL because there are fewer prefills slowing down decodes.
65: - Higher values achieve better time to first token (TTFT) as you can process more prefill tokens in a batch.
66: - For optimal throughput, we recommend setting `max_num_batched_tokens > 8192` especially for smaller models on large GPUs.
67: - If `max_num_batched_tokens` is the same as `max_model_len`, that's almost the equivalent to the V0 default scheduling policy (except that it still prioritizes decodes).

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • a universal default value
  • latency-optimal settings
  • released-version applicability without checking the target tag

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.