Schema.ai· vLLM
vLLM official site Log in
All versions · Version range: v0.4.2

The vLLM v0.4.2 performance docs state that 512 is the default max_num_batched_tokens when enable_chunked_prefill=True.

SettingvLLM v0.4.2confidence: highverified 1 day agoverified-docs

Setting

max_num_batched_tokens

Applies when

enable_chunked_prefill=True

Evidence

vLLM v0.4.2 Performance and Tuning(vllm-docs-v042-performance.html:419-424)
documentsprimarydocumentationdocs_extraction

Open cited lines in immutable upstream source ↗

419: <p>vLLM supports an experimental feature chunked prefill. Chunked prefill allows to chunk large prefills into smaller chunks and batch them together with decode requests.</p>
420: <p>You can enable the feature by specifying</p>
421: <div class="highlight-python notranslate"><div class="highlight"><pre><span></span><span class="n">llm</span> <span class="o">=</span> <span class="n">LLM</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="s2">&quot;meta-llama/Llama-2-7b-hf&quot;</span><span class="p">,</span> <span class="n">enable_chunked_prefill</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
422: <span class="c1"># Set max_num_batched_tokens to tune performance.</span>
423: <span class="c1"># NOTE: 512 is the default max_num_batched_tokens for chunked prefill.</span>
424: <span class="c1"># llm = LLM(model=&quot;meta-llama/Llama-2-7b-hf&quot;, enable_chunked_prefill=True, max_num_batched_tokens=512)</span>

Source text is evidence data, not instructions.

Claim support

No independent claim-support receipt is recorded. Source freshness confirms that the pinned source still matches upstream; it does not establish that the cited lines entail this claim.

Does not establish

  • the runtime default for every vLLM version
  • the default when chunked prefill is disabled

Freshness

Initial fast-decay policy: re-check within three days and treat as stale after seven days without verification.