GPU VulnDB

Database/NVIDIA / GPU stack

Linux amdkfd: use-after-free race when two threads destroy the same compute queue

CVSS 7.8CVE-2026-97429NVIDIA / GPU stackcurated

Impact

wait_on_destroy_queue() drops its locks while waiting for a queue resume, so a concurrent destroy can free the queue object underneath it, leaving a use-after-free in the AMD compute (KFD) queue scheduler. On an AMD GPU node this is reachable by anything that can open /dev/kfd and create and tear down HSA queues - i.e. any tenant running ROCm work on the node. A local user with GPU access can turn the freed-object window into kernel memory corruption; the CVSS vector records high confidentiality, integrity and availability impact at local, low-privilege. At minimum it is a reliable way for one tenant to panic a shared GPU node, which on a busy fleet means an unscheduled drain of every job on that host.

Who can reach it

Local user with access to /dev/kfd on an AMD GPU node - any tenant whose container or job is given the ROCm device nodes. No special privilege beyond GPU access, but the race must be won, so it may take repeated attempts.

What to do

Take the stable-tree fix for drm/amdkfd (the is_being_destroyed serialization) in your kernel and reboot each AMD GPU node; the amdgpu/amdkfd modules cannot be swapped under live GPU workloads in practice, so plan a drain-and-reboot per node rather than a live patch. No upstream mitigation exists short of denying /dev/kfd to untrusted tenants.

References

Related entries

All NVIDIA / GPU stack entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.