Linux drm/amdgpu: NULL dereference during GPU reset when KFD init failed after probe
Impact
amdgpu_amdkfd_clear_kfd_mapping() assumed a non-NULL kfd_dev had a fully populated node array. If KFD device initialization failed after probe - the upstream example is kgd2kfd_device_init() setting num_nodes and then bailing out on missing PCIe atomics support before allocating nodes[0], while the kfd_dev stays attached - a later GPU reset dereferences nodes[0]->id and panics the host. The exposure is narrow: it needs a node that already failed KFD bring-up, so the compute stack on that GPU was not working in the first place. What it turns into is a kernel oops at exactly the moment the driver is trying to recover a hung GPU, converting a recoverable reset into a lost node. Not an attacker-reachable boundary; a reliability bug on misconfigured or unsupported hardware.
Who can reach it
Local, and not meaningfully attacker-controlled: the trigger is a GPU reset on a device whose KFD initialization failed. No authentication or remote path involved.
What to do
Fixed in stable by requiring the authoritative KFD initialization flag before walking the node array; take the stable kernel carrying commit 157d3f1db7e7 / 7f9caa70aef0. Requires a kernel update and a reboot per node, so fold it into a normal drain-and-reboot cycle rather than scheduling an emergency window.
References
Related entries
- Linux kernel amdgpu: register BAR mapping leaks on every driver unload or hot-unplugCVE-2026-98179 · Linux kernel amdgpu (register BAR iounmap on device removal)Unscored
- amdgpu nbio_v7_9: NULL dereference in hard IRQ when a RAS interrupt arrives before RAS late_initCVE-2026-98270 · Linux kernel amdgpu nbio_v7_9 (RAS controller interrupt handler)Unscored
- GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not coveredNCVD-0000-012-gpu-accelerator-firmware-vbios-g · GPU / accelerator firmware (VBIOS, GSP, NVSwitch)Unscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-001-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-003-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent losslessNCVD-2023-003-integrated-gpu-graphics-data-com · Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.