Compute & power

Chips, clusters, energy, inference economics — who can run what, at what price.

Compute & powerThe Stack

SGLang ships day-zero MoE kernels faster than most teams pick a runtime

Apache-2.0 serving framework. v0.5.14 landed 26 June 2026 with DeepSeek-V4 and Waterfill/LPLB load balancing.

TLDR

SGLang is an open-source LLM and multimodal serving runtime from the sgl-project, hosted under LMSYS. Release v0.5.14 on 26 June 2026 added GLM-5.2, Kimi-K2.7-Code, Waterfill/LPLB MoE dispatch, and DeepSeek-V4 NVFP4 paths documented in release notes. LICENSE is Apache-2.0. GitHub star count at publication is UNKNOWN.

Compute & powerDispatch

NVIDIA says Vera Rubin is in full production

The 31 May GTC Taipei post puts shipments in “this fall,” names Dell, HPE, Lenovo, and Supermicro, and still says products are when-and-if-available.

TLDR

NVIDIA's 31 May 2026 newsroom post, from GTC Taipei, says the Vera Rubin platform is ramping into full production with partners across 350-plus factories in 30 countries, including Dell, HPE, Lenovo, and Supermicro. It claims 10x agent throughput at scale versus the Grace Blackwell platform. Production shipments "are set to begin starting this fall." The same page's forward-looking and when-and-if-available disclaimers mean a purchase order is not a delivery date. Independent throughput measurements are UNKNOWN.

Compute & powerThe Record

DeepSeek-V4 paper reports 27% FLOPs and 10% KV cache at 1M tokens versus V3.2

arXiv:2606.19348 (April 26, 2026); MIT checkpoints on HuggingFace; CSA/HCA hybrid attention.

TLDR

DeepSeek-V4-Pro and V4-Flash use interleaved Compressed Sparse Attention and Heavily Compressed Attention for 1M-token contexts. At 1M tokens the paper reports V4-Pro at 27% single-token FLOPs and 10% KV cache versus DeepSeek-V3.2; V4-Flash at ~10% FLOPs and ~7% KV. MIT weights on HuggingFace. Dollar cost per million agent tokens on your stack is UNKNOWN until benchmarked.

Compute & powerDispatch

H200 and MI325X China licenses go case-by-case, still not a general license

BIS final rule 2026-00789, effective 15 January 2026, reviews U.S. exports of those SKUs case by case if exporters certify U.S. supply, a 50 percent TPP cap, KYC, and U.S. third-party testing. Reexports stay denied.

TLDR

BIS published final rule 2026-00789 on 15 January 2026, effective the same day, changing license review for certain advanced computing commodities, named examples NVIDIA H200 and AMD MI325X, from presumption of denial to case-by-case for exports from the United States to end-users in China or Macau. Applicants must certify U.S. commercial availability, no delay to U.S. orders, no foundry diversion, aggregate TPP to China and Macau no more than 50 percent of U.S. end-use shipments of the same product, KYC and remote-access screening, and U.S. third-party testing. Reexports, in-country transfers, and shipments to D:5-headquartered entities remain presumption of denial. How many licenses BIS will grant is UNKNOWN.