GPU VulnDB

Database/NVIDIA / GPU stack

Linux drm/amdgpu: NULL dereference during GPU reset when KFD init failed after probe

UnscoredCVE-2026-98178NVIDIA / GPU stackcurated

Impact

amdgpu_amdkfd_clear_kfd_mapping() assumed a non-NULL kfd_dev had a fully populated node array. If KFD device initialization failed after probe - the upstream example is kgd2kfd_device_init() setting num_nodes and then bailing out on missing PCIe atomics support before allocating nodes[0], while the kfd_dev stays attached - a later GPU reset dereferences nodes[0]->id and panics the host. The exposure is narrow: it needs a node that already failed KFD bring-up, so the compute stack on that GPU was not working in the first place. What it turns into is a kernel oops at exactly the moment the driver is trying to recover a hung GPU, converting a recoverable reset into a lost node. Not an attacker-reachable boundary; a reliability bug on misconfigured or unsupported hardware.

Who can reach it

Local, and not meaningfully attacker-controlled: the trigger is a GPU reset on a device whose KFD initialization failed. No authentication or remote path involved.

What to do

Fixed in stable by requiring the authoritative KFD initialization flag before walking the node array; take the stable kernel carrying commit 157d3f1db7e7 / 7f9caa70aef0. Requires a kernel update and a reboot per node, so fold it into a normal drain-and-reboot cycle rather than scheduling an emergency window.

References

Related entries

All NVIDIA / GPU stack entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.