Database/Kernel, userspace & hypervisor
Linux kernel nvme-fabrics: admin tag exhaustion during reconnect can hang the host indefinitely
Impact
NVMe-over-Fabrics register reads and writes used ordinary admin queue tags. A local user issuing many NVMe admin commands at once can exhaust admin_q tags; if a controller reset or an I/O timeout lands while those commands are outstanding, the reconnect path cannot get a tag to update the controller registers and the kernel hangs forever rather than recovering. Rated 5.5, local, availability-only. On a GPU fleet this bites where NVMe-oF (NVMe/TCP or NVMe/RDMA) carries dataset and checkpoint storage: a wedged host does not come back on its own, so a node mid-training stops making progress and has to be power-cycled, taking its accelerators out of the pool and losing whatever was not checkpointed.
Who can reach it
Local user or workload able to submit NVMe admin commands to a fabrics-attached controller (for example via the nvme CLI on an ioctl-accessible device). Authentication to the host is required; the trigger also occurs accidentally under heavy admin traffic plus a reset or I/O timeout.
What to do
Move to a stable kernel carrying the fix, which makes reg_read32/reg_read64/reg_write32 use reserved tags so the reconnect path always has one available; the stable commits are listed in the references. Deployment is a node reboot - drain the GPU node first. Until then, restrict who can issue NVMe admin ioctls on fabrics controllers. No vendor advisory or packaged fixed version appears in the record.
References
Related entries
- Linux drm/xe GPU kernel driver (display opregion): A resource leak in the xe driver's display opregion handlingCVE-2024-44980 · Linux drm/xe GPU kernel driver (display opregion)Medium
- Linux kernel (drivers/pci/hotplug): The hotplug driver disables MSI/MSI-X during slot unregistration after the MSI dataCVE-2024-46761 · Linux kernel (drivers/pci/hotplug)Medium
- Linux kernel (drivers/gpu/drm/xe): Same per-client accounting path, different failure - if the fdinfo read drops theCVE-2024-46867 · Linux kernel (drivers/gpu/drm/xe)Medium
- Linux kernel BPF verifier: sign-extended packet-pointer loads produce an invalid skb->data addressCVE-2024-47702 · Linux kernel BPF verifier (sign-extended loads of __sk_buff data/data_end/data_meta)Medium
- Linux kernel (drivers/gpu/drm/xe): User VM_BIND work is scheduled onto engines that can themselves take page faultsCVE-2024-47729 · Linux kernel (drivers/gpu/drm/xe)Medium
- Linux kernel x86/mm: identity maps built from 1 GB pages cover unrequested reserved memoryCVE-2024-50017 · Linux kernel x86 identity mapping (ident_pud_init GB-page mapping)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.