GPU VulnDB

Database/NVIDIA / GPU stack

Linux kernel habanalabs: early mem_mgr IDR teardown leaks memory buffers held by a live Gaudi context

CVSS 5.5CVE-2023-53353NVIDIA / GPU stackcurated

Impact

The habanalabs driver destroyed its memory-manager IDR when the user closed the device file descriptor, while the user context and its buffers could still be in use. Later attempts to release those buffers fail because their handles are no longer in the IDR, and the memory is never freed. The result is an unbounded kernel memory leak driven by ordinary open/close cycles of the accelerator device, rated 5.5 with availability-only impact. On a Gaudi node this is a slow availability problem rather than a privilege issue: repeated training or inference jobs that open /dev/accel and exit leak kernel memory until the node degrades or OOMs, and recovering it requires rebooting a node whose accelerators are otherwise healthy. A tenant that can open the accelerator device can drive the leak deliberately.

Who can reach it

Local user or container holding access to the habanalabs accelerator character device (/dev/accel/accel*) - that is, any tenant with a Gaudi device mapped into its pod. Authentication to the host is required; no remote path.

What to do

Update to a stable kernel containing the fix, which splits IDR destruction out of memory-manager fini and defers it to hpriv_release() once no user context remains and no buffers are in use; the stable commits are in the references. Rolling it out means draining the Gaudi node and rebooting it. As an interim measure, monitor kernel slab growth on accelerator nodes and recycle them on a schedule. The record names no vendor advisory or packaged fixed version.

References

Related entries

All NVIDIA / GPU stack entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.