Schema.ai· vLLM
vLLM official site ↗vLLM v0.23.0API up

vLLM operational knowledgeSCHEMA.AI

vLLM (official site) is an open-source inference and serving engine for large language models. Schema.ai curates operational knowledge for running it — source-grounded, version-scoped, evidence-backed building blocks, not generated answers.

Search results for "number of tokens queried and the number of queried tokens"

Focusedmatch: mediumquestionvLLM v0.23.0

Source-grounded blocks matched your query.

Answerability

answerable_from_corpus

The local vLLM corpus has source-grounded blocks matching the query terms.

1 matching claim

Each result is a source-grounded claim — click through for its evidence, guards, and graph context. Ranked by: Lexical overlap against claim text, object values, applicability, and entity names.

1Claimcaveatconfidence: highfreshness: fresh · full

In vLLM v0.23.0, vllm:prefix_cache_queries is a Prometheus counter measured in queried tokens it counts tokens, not requests or cache blocks.

Applies when
vllmv0.23.0PrometheusStatLogger metrics publisherV1 metricsbasis: vLLM v0.23.0 source-code metric registration, scheduler stats, and metrics design documentation
Topics
Sources
source.vllm.code.v1_metrics_loggers.v0230source.vllm.code.v1_metrics_loggers.v0230
Does not establish (3)
  • a healthy hit-rate threshold
  • request-level cache behavior (the unit is tokens, not requests)
  • cache sizing guidance

Matched query terms against entity.vllm.metric.prefix_cache_queries and v0.23.0.

View claim →

Provenance & sources

Scope: local_corpus_snapshots

vllm-projectsource_filev0.23.0

How to use this

Use as

Source-grounded building blocks for a downstream UI, evaluator, or human.

Do not use as

A complete recommendation or comprehensive vLLM coverage map.

Next steps
  • Inspect returned blocks and their source snapshots.Schema returns evidence-backed ingredients rather than a composed answer.