GPU VulnDB

Database/Kernel, userspace & hypervisor

Linux kernel nvme-rdma: double cleanup and DMA unmap after request completion on the -EIO path

CVSS 7.0CVE-2026-98154Kernel, userspace & hypervisorcurated

Impact

On the -EIO path in the NVMe-over-RDMA submission routine, the host path error helper completes the request and the code then still cleans up the command and unmaps the SQE DMA - a use-after-completion with a DMA unmap on an already-completed request. This reaches GPU nodes that mount NVMe-oF over RDMA for datasets and checkpoints, which is common where the fabric is InfiniBand or RoCE. Triggering it needs the error path to be hit, so the realistic route is fabric instability or a misbehaving target rather than a tenant action; the likely outcome is a host crash and the loss of every job on the node.

Who can reach it

Not a tenant-facing path. Requires the nvme-rdma error path to fire, which follows from storage fabric or target-side faults; anyone who can disrupt the storage fabric or control an NVMe-oF target the host connects to is in a position to provoke it.

What to do

Apply the stable kernel fix and reboot each affected node. Only hosts that actually use NVMe over RDMA are affected - if your GPU nodes mount storage another way, this can ride along with the next scheduled kernel window rather than driving one.

References

Related entries

All Kernel, userspace & hypervisor entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.