A wrong-results bug in CUDA 13.4's cuBLAS
NVIDIA has shipped cuBLAS 13.8.1, a standalone patch for a bug in the math library that came with CUDA 13.4. The cuBLAS patch release notes say cuBLASLt Grouped GEMM operations "could produce incorrect results on Blackwell and Rubin GPUs". The notes trace the bug to cuBLAS 13.7.0, the version in CUDA 13.4. They name no workaround and no affected shapes or data types.
The Frontier covered the CUDA 13.4 release itself in September. See CUDA 13.4: Windows on Arm, Rubin sm_107 preview, MPS V3, and a CDMM default flip.
NVIDIA's cuBLAS team, outside the toolkit cycle
The patch notes page lists cuBLAS releases "issued separately from CUDA Toolkit major and minor releases". Patch 13.8.1 applies to CUDA Toolkit 13.4 Update 1. NVIDIA's Ubuntu 24.04 package repository lists the 13.8.1 packages from 26 September. The nvidia-cublas 13.8.1.7 wheel reached PyPI on 28 September UTC, which was 29 September in Singapore.
The cuBLAS 13.8.1 download page offered Linux builds only, for x86_64 and arm64-sbsa, when The Frontier checked at 02:43 SGT on 8 October. The PyPI wheels are Linux-only too.
Update 1 listed the bug and kept the build
Three NVIDIA documents show how the bug surfaced:
- The CUDA 13.4 release notes, now in NVIDIA's archive, do not list it.
- The CUDA 13.4 Update 1 release notes, last updated 21 September, list it as a known cuBLAS issue.
- The patch notes mark it fixed in 13.8.1, with NVIDIA bug number 6798616.
Update 1 ships cuBLAS 13.8.0.4, so it carries the bug. The same update fixes four other cuBLAS bugs that could return incorrect results, two of them on B300 and Rubin GPUs with NVFP4 inputs. It adds support for PTX ISA 9.4. NVIDIA says it raises emulated FP64 ZGEMM peak performance by up to 25% on Rubin. Its cuSOLVER update also fixes a hang inside CUDA green contexts, which NVIDIA explained in a 6 October blog post.
The Frontier counted the cuBLAS sections of the CUDA 13.x release notes. They describe four separate Grouped GEMM issues that could return incorrect results. One dates from 13.1, one from 13.2 Update 1 and two from 13.4. CUDA 13.4 also added dynamic scheduling that NVIDIA says speeds Grouped GEMM calls with many groups by up to 20% on Blackwell data center GPUs. The notes do not say what caused the new bug. The cuBLAS documentation labels its grouped matrix layout attributes experimental.
Pip pins decide who still runs the bug
Package metadata on PyPI decides which cuBLAS many Python users get:
- The newest cuda-toolkit metapackage, 13.4.2, pins nvidia-cublas 13.8.0.4, the build with the known issue.
- A pip environment built from that metapackage keeps the affected build unless someone overrides the pin.
- PyTorch 2.14.1 pins cuda-toolkit 13.0.3, which pins nvidia-cublas 13.1.1.3. That build is older than cuBLAS 13.7.0, where the bug began.
- The JAX CUDA 13 plugin 0.11.2 sets only a floor of nvidia-cublas 13.0.0.19.
- From late on 9 September to 28 September UTC, the newest nvidia-cublas on PyPI was 13.7.0.27 or 13.8.0.4, both affected versions.
- A fresh JAX install in that window would have pulled an affected build. The Frontier has not checked whether JAX calls the Grouped GEMM path.
NVIDIA's notes are the only public account of the bug, the fix and the speed figures. The Frontier tested none of them, and a GitHub issue search at 02:42 SGT found no public report citing bug 6798616.
No new driver branch came with the patch. NVIDIA's data center driver documentation still lists R615, version 615.71.09 on Linux, as its newest branch. The Frontier's earlier watchlist covers the 13.4 split between driver and toolkit installs. See CUDA 13.4 upgrade watchlist: CDMM default, driver and toolkit split, MPS V3.
Check cuBLAS versions on Blackwell and Rubin nodes
- Teams on CUDA 13.4 or 13.4 Update 1 with Blackwell or Rubin GPUs should check whether their code or libraries call cuBLASLt Grouped GEMM.
- NVIDIA's notes name cuBLAS 13.7.0 and 13.8.0 as the affected versions and 13.8.1 as the fix.
- In pip environments,
pip show nvidia-cublasprints the installed wheel version. - Teams that install cuda-toolkit 13.4.2 through pip need to install nvidia-cublas 13.8.1.7 over the pinned version.
- Windows teams should watch the download page, which listed no Windows build at the check.
NVIDIA's repository also shows other CUDA-X releases since 13.4, including cuDNN 9.27.0 on 29 September and NCCL 2.32.3 from 18 September.
