Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx8): A correctness defect in the amdgpu RAS / GPU
Impact
A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx8: drop unecessary BUG_ON()
Who can reach it
Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.
What to do
Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.
References
Related entries
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)CVE-2026-68436 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)Unscored
- Linux kernel drm/xe: memory leak when hang replay state is set twice on an exec queueCVE-2026-74699 · Linux kernel drm/xe (exec_queue_set_hang_replay_state)Unscored
- Linux kernel AMD XDNA driver: unprivileged mmap plus MADV_DONTNEED trips a BUG_ONCVE-2026-74716 · Linux kernel accel/amdxdna (AMD XDNA NPU driver, amdxdna_insert_pages)Unscored
- Linux kernel AMD XDNA driver: error path closes a live VMA and underflows its referencesCVE-2026-74721 · Linux kernel accel/amdxdna (amdxdna_insert_pages error paths call vm_ops->close)Unscored
- Linux kernel amdgpu: duplicate FENCE chunks in one submission leak a buffer-object reference per submitCVE-2026-80539 · Linux kernel amdgpu (amdgpu_cs_pass1, duplicate AMDGPU_CHUNK_ID_FENCE chunks)Unscored
- Linux kernel amdgpu UVD: decode image size computed from width instead of pitchCVE-2026-80540 · Linux kernel amdgpu UVD (decode image minimum size validation, unbounded pitch)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.