Linux kernel amdgpu: use of freed module text when RCU callbacks run after module unload
Impact
amdgpu did not call rcu_barrier() in its module exit path, so a pending call_rcu() callback could fire after the module text was freed, jumping into unmapped memory and panicking the kernel - the reported trace is a page fault in rcu_do_batch from softirq context. This only bites on amdgpu module unload, which on a GPU node means driver reinstall or upgrade, a ROCm stack update, or an automated driver rollout; the failure mode is a host crash that takes every workload on the node with it and needs a reboot. It is not a tenant-reachable attack path: unloading the module requires root on the host. Treat it as a stability fix to fold into the next driver-upgrade window rather than an urgent security patch.
Who can reach it
Local root on the host, since only a privileged user can unload the amdgpu module. Not reachable by tenants holding /dev/kfd or /dev/dri, and not reachable over the network.
What to do
Pick up the stable kernel containing commit 67a654b41cfa (upstream feaa5039f6c1) from your distribution's kernel updates. Applying it means installing a new kernel and rebooting the node, so drain the GPU node and fold it into a scheduled reboot window - there is no runtime mitigation short of not unloading amdgpu, which is itself a reasonable interim rule: reboot rather than rmmod when updating the driver.
References
Related entries
- GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not coveredNCVD-0000-012-gpu-accelerator-firmware-vbios-g · GPU / accelerator firmware (VBIOS, GSP, NVSwitch)Unscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-001-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-003-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent losslessNCVD-2023-003-integrated-gpu-graphics-data-com · Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)Unscored
- NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode: NVIDIA GPU confidentialNCVD-2023-004-nvidia-confidential-computing-h1 · NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating modeUnscored
- Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent losslessNCVD-2023-005-integrated-gpu-graphics-data-com · Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.