Linux kernel amdkfd: unbounded copy_to_user in get_wave_state leaks adjacent kernel memory
Impact
get_wave_state() for v9 trusts the control-stack size and offset read from the MQD. Those fields are attacker-controlled on the CRIU restore path via AMDKFD_IOC_RESTORE_PROCESS, and an offset larger than the size underflows into a roughly 4 GiB read length; a 1 MiB size against a 4 KiB allocation leaks 1 MiB of adjacent kernel memory. What leaks is exactly the sensitive neighbourhood on an AMD GPU node: other queues' MQDs, ring buffers, and KASLR pointers. On a shared AMD fleet this is a local tenant reading kernel memory that belongs to other GPU workloads, and it is a plausible first step toward defeating KASLR for a further exploit.
Who can reach it
Local user holding /dev/kfd - any tenant with an AMD GPU pod or shell on the node. No special privilege beyond device access, and the CRIU restore ioctl is reachable from an ordinary process.
What to do
Apply the stable kernel fix (commits 7ef1444, d9183d9, ec64668) that clamps both the control-stack size and offset to the allocated buffer, then reboot each AMD GPU node - drain and reboot, since a live kernel patch is not offered here. If you cannot reboot soon, the exposure is confined to processes holding /dev/kfd, so restricting which workloads get the device is a real mitigation.
References
Related entries
- NVIDIA vGPU Manager (vGPU plugin): Stack buffer overflow in the vGPU Manager with enough control for a guest to place aCVE-2021-1099 · NVIDIA vGPU Manager (vGPU plugin)High
- NVIDIA vGPU Manager (vGPU plugin): A guest-supplied string may not be null-terminated, and the host plugin reads pastCVE-2021-1120 · NVIDIA vGPU Manager (vGPU plugin)High
- NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The vGPU plugin double-frees hostCVE-2022-31614 · NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)High
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)CVE-2022-50079 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)High
- NVIDIA/Mellanox ConnectX driver (mlx5_ib UMR resource init/cleanup): A failed workqueue allocation during mkey-cacheCVE-2023-52851 · NVIDIA/Mellanox ConnectX driver (mlx5_ib UMR resource init/cleanup)High
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)CVE-2024-26660 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.