Database/Kernel, userspace & hypervisor
Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute
Impact
A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: amd/amdkfd: resolve a race in amdgpu_amdkfd_device_fini_sw
Who can reach it
Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.
What to do
Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.
References
Related entries
- Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A division by zero in the amdkfd (KFD compute driverCVE-2025-68174 · Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd)High
- Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFDCVE-2025-40332 · Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)Unscored
- Xen / x86 CPU: Floating Point Divider State Sampling - transient-execution leak of FP divider state across domainsCVE-2025-54505 · Xen / x86 CPUUnscored
- Xen (Viridian): Incorrect input sanitisation in Viridian (Hyper-V enlightenment) hypercallsCVE-2025-58147 · Xen (Viridian)Unscored
- Linux kernel bpf: BPF_REFCOUNT field was not marked unique in the verifier's field checksCVE-2026-100074 · Linux kernel BPF verifier (bpf_refcount not marked as a unique field)Unscored
- Xen (EPT): Use-after-free of EPT paging structures - HVM guest to host compromiseCVE-2026-23554 · Xen (EPT)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.