Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): An out-of-bounds access in the amdgpu RAS / GPU
Impact
An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpu_coredump ring dump
Who can reach it
Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.
What to do
Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.
References
Related entries
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPUCVE-2023-52585 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPUCVE-2023-52814 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A race condition or locking defect in the amdgpuCVE-2023-53074 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): An out-of-bounds access in the amdgpu RAS / GPUCVE-2024-26915 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPUCVE-2024-43908 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
- Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU resetCVE-2024-44961 · Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.