Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPU
Impact
A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/ras: Fix NULL deref in ras_core_ras_interrupt_detected()
Who can reach it
Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.
What to do
Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.
References
Related entries
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPUCVE-2026-53315 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras)Medium
- NVIDIA License System - DLS virtual appliance: Installation scripts on the DLS appliance leave other users' credentialsCVE-2022-21818 · NVIDIA License System - DLS virtual applianceMedium
- Triton Inference Server: Insufficient access-control granularityCVE-2024-0103 · Triton Inference ServerMedium
- Intel Gaudi / gaudi-container-runtime: A path-traversal bug in the container runtime shim that wires Gaudi devices intoCVE-2026-32677 · Intel Gaudi / gaudi-container-runtimeMedium
- NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Improper access control in the escape handler leaks informationCVE-2021-1055 · NVIDIA Windows GPU Display Driver (nvlddmkm.sys)Medium
- NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An off-by-one error in nvidia.ko permits data tamperingCVE-2022-34684 · NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.