GPU VulnDB

Database/NVIDIA / GPU stack

Linux amdgpu userq: MQD and firmware buffer objects can be evicted, hanging the GPU node

UnscoredCVE-2026-97498NVIDIA / GPU stackcurated

Impact

User queues need their MQD and firmware buffer objects mapped for the lifetime of the queue, but they were not pinned, so memory pressure could evict them while hardware still held the old addresses. The result is GPU page faults and, per the commit message, a system hang. On a GPU node that means the whole host stops serving - every tenant on the box, not just the one that caused the eviction - and recovery is a reboot, which is the expensive kind of outage on an accelerator fleet. This is availability only; the record describes no memory disclosure or privilege gain and carries no CVSS score.

Who can reach it

Local user of the amdgpu user-queue interface: any tenant with a GPU device that can create user queues and then drive enough VRAM pressure to trigger eviction. No authentication beyond having the GPU.

What to do

Move to a stable kernel with the MQD/firmware BO pinning fix (two stable commits listed) and reboot the node. No vendor advisory or packaged fixed version appears in the record. Until then the only lever is keeping GPU memory pressure away from user-queue workloads, which is not a real mitigation.

References

Related entries

All NVIDIA / GPU stack entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.