<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Inference Lab</title><link>https://inference-lab-engineering.xyl5869.chatgpt.site</link><description>Practical guides, benchmarks, and troubleshooting for LLM inference, GPUs, vLLM, SGLang and production AI infrastructure.</description><item><title>Understanding FP8 KV Cache in vLLM</title><description>Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/fp8-kv-cache-vllm</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/fp8-kv-cache-vllm</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Prefix Cache vs KV Cache in LLM Inference</title><description>Distinguish the attention state stored during generation from reuse of that state across requests.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/prefix-cache-vs-kv-cache</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/prefix-cache-vs-kv-cache</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Qwen3.8-27B on 4×RTX 4090: Deployment Guide</title><description>A deployment validation plan for Qwen3.8-27B, covering hardware inventory, version pinning and a reproducible acceptance test.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/qwen3-8-27b-4x4090-sglang</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/qwen3-8-27b-4x4090-sglang</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>SGLang MTP / NEXTN Explained</title><description>Separate speculative decoding concepts from model- and release-specific launch options.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/sglang-mtp-nextn</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/sglang-mtp-nextn</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>vLLM vs SGLang: Choosing an Inference Engine</title><description>Choose an inference engine using workload compatibility, reproducibility and operational requirements—not an unqualified leaderboard.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/vllm-vs-sglang</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/guides/vllm-vs-sglang</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item><item><title>NVIDIA Xid 79: GPU Has Fallen Off the Bus</title><description>Preserve evidence, isolate the affected GPU and investigate PCIe accessibility without assuming a single root cause.</description><link>https://inference-lab-engineering.xyl5869.chatgpt.site/troubleshooting/nvidia-xid-79-gpu-fallen-off-bus</link><guid>https://inference-lab-engineering.xyl5869.chatgpt.site/troubleshooting/nvidia-xid-79-gpu-fallen-off-bus</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate></item></channel></rss>