Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amd/amdgpu/amdgpu_cs): A memory or reference-count
Impact
A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/amdgpu/amdgpu_cs: fix refcount leak of a dma_fence obj
Who can reach it
Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.
What to do
Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.
References
Related entries
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display coreCVE-2022-49232 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)Medium
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display coreCVE-2022-49233 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)Medium
- Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)CVE-2022-49294 · Linux kernel amdgpu display core (DC/DM) (drm/amd/display)Medium
- Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs): Missing or insufficient validation ofCVE-2022-49335 · Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs)Medium
- Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): An out-of-bounds access in the amdgpu display core (DC/DM)CVE-2022-49365 · Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)Medium
- Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu powerCVE-2022-49529 · Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.