GPU VulnDB

Database/NVIDIA / GPU stack

Linux kernel amdkfd: user-controlled metadata size lets any render-group user force a huge kernel allocation

UnscoredCVE-2026-93823NVIDIA / GPU stackcurated

Impact

The AMDKFD_IOC_GET_DMABUF_INFO ioctl allocated the buffer for returning buffer-object metadata using a size supplied entirely by user space. Any process able to open the ROCm compute device and issue the ioctl could request an order-MAX allocation (the commit log cites 2 GiB) and drive the kernel into OOM in kernel context. On a shared AMD GPU node that is a single-tenant denial of service against the whole host: every other training or inference job on the node is collateral, and the node has to be drained and rebooted rather than just having one pod killed. The fix makes the driver determine the real metadata size itself and ask user space to retry with a correctly sized buffer instead of trusting the request.

Who can reach it

Local user with access to the AMD KFD compute device - in practice any tenant holding /dev/kfd, i.e. any container that has been given an AMD GPU. No special privilege beyond render-group membership, and no authentication beyond having a GPU pod on the node.

What to do

Pick up the fix from the stable trees linked in the record (four backports) and run a patched kernel. This is a kernel driver change, so it means rebooting each AMD GPU node after draining its workloads; there is no module-reload path that is safe while jobs hold KFD contexts. Until then, the exposure follows GPU device access - tenants that do not get /dev/kfd cannot reach the ioctl.

References

Related entries

All NVIDIA / GPU stack entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.