The Frontier · Tags
#inference
2 published articles
SGLang ships day-zero MoE kernels faster than most teams pick a runtime
Apache-2.0 serving framework. v0.5.14 landed 26 June 2026 with DeepSeek-V4 and Waterfill/LPLB load balancing.
DeepSeek-V4 paper reports 27% FLOPs and 10% KV cache at 1M tokens versus V3.2
arXiv:2606.19348 (April 26, 2026); MIT checkpoints on HuggingFace; CSA/HCA hybrid attention.