Linux drm/ttm: swapped-out resources stay in their bulk_move range, leaving a dangling cursor (use-after-free)
Impact
A regression in TTM's swapout path meant the bulk_move bookkeeping was skipped on every successful swapout, so a swapped-out GPU resource stayed inside its buffer object's bulk_move range. When that resource is later freed or the BO leaves the range, a range endpoint is left pointing at freed memory and the next bulk-move operation is a use-after-free - observed as list corruption or a NULL dereference minutes to hours later, at process exit or reboot. On a GPU node this is a memory-corruption crash in the graphics memory manager shared by amdgpu, so it is a stability and potentially exploitable-corruption issue in a kernel path any local GPU workload can drive through memory pressure. The reporter reproduced it through hibernation cycles on a consumer APU; datacenter nodes do not usually hibernate, which makes the swapout pressure needed to reach it less common but not absent.
Who can reach it
Local only. A user with access to a GPU device node whose workload drives TTM into swapping out buffer objects; no network path and no authentication boundary is crossed. Reachable only on hosts running a TTM-based GPU driver, which on a GPU fleet means amdgpu.
What to do
Fixed in stable by testing for success correctly in ttm_bo_swapout_cb(); take the stable kernel carrying commit 1169fe8c11ca / 3db7d7d583419. Applying it means a kernel update and a reboot of each node, so drain the node first - the driver module cannot be reloaded safely under live GPU workloads.
References
Related entries
- Linux drm/amdkfd: integer underflow in the EOP ring size log calculationCVE-2026-98176 · Linux kernel drm/amdkfd (EOP ring size field in cp_hqd_eop_control)Unscored
- Linux drm/amdgpu: NULL dereference during GPU reset when KFD init failed after probeCVE-2026-98178 · Linux kernel drm/amdgpu (amdgpu_amdkfd_clear_kfd_mapping on GPU reset)Unscored
- Linux kernel amdgpu: register BAR mapping leaks on every driver unload or hot-unplugCVE-2026-98179 · Linux kernel amdgpu (register BAR iounmap on device removal)Unscored
- amdgpu nbio_v7_9: NULL dereference in hard IRQ when a RAS interrupt arrives before RAS late_initCVE-2026-98270 · Linux kernel amdgpu nbio_v7_9 (RAS controller interrupt handler)Unscored
- GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not coveredNCVD-0000-012-gpu-accelerator-firmware-vbios-g · GPU / accelerator firmware (VBIOS, GSP, NVSwitch)Unscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-001-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.