NVIDIA's top-end consumer GPU, the GeForce RTX 5090, and top-of-the-line ProViz SKU, the RTX 6000 PRO, have been plagued by a new virtualization bug. According to developers at CloudRift, who are creating a GPU cloud for AI developers, they have encountered a specific bug that renders the RTX 5090 and RTX PRO 6000 completely unresponsive. After a few days or weeks of constant usage, the GPU virtual machine can completely freeze without a sign of responsiveness. This occurs at random times, with no clear indication of why. The team has tested multiple GPUs, including the H100, B200, and older RTX 4090 models, all of which showed no issues. Not even the highest-performing server-grade B200 GPU from the "Blackwell" family experiences these issues, but consumer and ProViz SKUs do.
Behind the scenes, a more technical approach explains the process of locking the GPU. When a GPU is handed off to a virtual machine via KVM and VFIO, the host performs a PCIe function-level reset (FLR) as part of the normal cleanup process when the VM stops or the device is moved. Instead of coming back online after that reset, the card becomes unresponsive. The kernel times out and reports the failure with the message "not ready 65535ms after FLR; giving up." Hence, the only point of failure is the GPU itself, and CloudRift has even issued a $1,000 bug bounty for anyone who can resolve the issue.