SGLang1 source
SGLang v0.5.12 released on 2026-05-16, featuring full inference support for DeepSeek V4, including various parallelism (tensor, expert, context, data parallel attention), hardware support (Nvidia B300/B200/H200/H100/GB200/GB300, AMD MI35X), prefill-decode disaggregation, sparse KV cache offloading (HiSparse), reasoning and tool call parsers, custom kernels (DeepGemm, FlashMLA, MegaMoE), and post-day-0 additions: HiCache under unified Radix Tree, W4A4 and W4A8 MoE kernels, compression kernels, TP16 support, fused quantization kernel, optimized MHC+DeepGemm pipeline, non-standard chat template support, multi-detokenizer support, pipeline parallelism + PD support. Also includes a unified Docker tag lmsysorg/sglang.
Why it mattersThis development may affect product decisions, technical choices or the direction of the AI market. Only one public source is currently available.
Sources & timeline