disable_hybrid_kv_cache_manager
SchedulerConfig flag forcing the KV cache manager to allocate the same KV cache size for all attention layer types.
SettingvLLM v0.23.0verified Jul 9
Claims (0)
Availability
disable_hybrid_kv_cache_manager is exposed on SchedulerConfig in the v0.23.0 source snapshot.vllm.config.scheduler_configcomplete_inventorySchedulerConfig dataclass field surface in vllm/config/scheduler.py