GPU VulnDB

Database/Kernel, userspace & hypervisor

Linux kernel nvme-fabrics: admin tag exhaustion during reconnect can hang the host indefinitely

CVSS 5.5CVE-2024-41082Kernel, userspace & hypervisorcurated

Impact

NVMe-over-Fabrics register reads and writes used ordinary admin queue tags. A local user issuing many NVMe admin commands at once can exhaust admin_q tags; if a controller reset or an I/O timeout lands while those commands are outstanding, the reconnect path cannot get a tag to update the controller registers and the kernel hangs forever rather than recovering. Rated 5.5, local, availability-only. On a GPU fleet this bites where NVMe-oF (NVMe/TCP or NVMe/RDMA) carries dataset and checkpoint storage: a wedged host does not come back on its own, so a node mid-training stops making progress and has to be power-cycled, taking its accelerators out of the pool and losing whatever was not checkpointed.

Who can reach it

Local user or workload able to submit NVMe admin commands to a fabrics-attached controller (for example via the nvme CLI on an ioctl-accessible device). Authentication to the host is required; the trigger also occurs accidentally under heavy admin traffic plus a reset or I/O timeout.

What to do

Move to a stable kernel carrying the fix, which makes reg_read32/reg_read64/reg_write32 use reserved tags so the reconnect path always has one available; the stable commits are listed in the references. Deployment is a node reboot - drain the GPU node first. Until then, restrict who can issue NVMe admin ioctls on fabrics controllers. No vendor advisory or packaged fixed version appears in the record.

References

Related entries

All Kernel, userspace & hypervisor entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.