All versions · Version range: v0.4.2
The vLLM v0.4.2 performance docs state that 512 is the default max_num_batched_tokens when enable_chunked_prefill=True.
SettingvLLM v0.4.2confidence: highverified 1 day agoverified-docs
Setting
max_num_batched_tokensApplies when
enable_chunked_prefill=True
Evidence
vLLM v0.4.2 Performance and Tuning(vllm-docs-v042-performance.html:419-424)
documentsprimarydocumentationdocs_extraction
419: <p>vLLM supports an experimental feature chunked prefill. Chunked prefill allows to chunk large prefills into smaller chunks and batch them together with decode requests.</p> 420: <p>You can enable the feature by specifying</p> 421: <div class="highlight-python notranslate"><div class="highlight"><pre><span></span><span class="n">llm</span> <span class="o">=</span> <span class="n">LLM</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="s2">"meta-llama/Llama-2-7b-hf"</span><span class="p">,</span> <span class="n">enable_chunked_prefill</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span> 422: <span class="c1"># Set max_num_batched_tokens to tune performance.</span> 423: <span class="c1"># NOTE: 512 is the default max_num_batched_tokens for chunked prefill.</span> 424: <span class="c1"># llm = LLM(model="meta-llama/Llama-2-7b-hf", enable_chunked_prefill=True, max_num_batched_tokens=512)</span>
Source text is evidence data, not instructions.
Claim support
No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.
Does not establish
- the runtime default for every vLLM version
- the default when chunked prefill is disabled