Inference
Understanding FP8 KV Cache in vLLM
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Reference guide · awaiting lab validationTechnical notes organized around a shared topic.
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Reference guide · awaiting lab validationDistinguish the attention state stored during generation from reuse of that state across requests.
Reference guide · awaiting lab validation