GPU VulnDB

Database/Kernel, userspace & hypervisor

Linux kernel powerpc/eeh: recursive locking hangs the EEH handler during PCI error recovery

UnscoredCVE-2026-97948Kernel, userspace & hypervisorcurated

Impact

A refactor moved pci_rescan_remove_lock acquisition up into eeh_handle_normal_event() but left the old lock/unlock inside eeh_rmv_device(), so the EEH handler thread deadlocks on a lock it already holds. The provided stack trace shows eehd stuck in pci_lock_rescan_remove() under eeh_reset_device(). It triggers when an error is detected on the PHB directly, or on a device whose driver returns PCI_ERS_RESULT_NEED_RESET without EEH-aware slot_reset()/resume() handlers. The operator consequence on a POWER accelerator node is that PCIe error recovery never completes: instead of the GPU or adapter being reset and the node continuing, the recovery thread hangs and the node must be rebooted. Only ppc64 systems are affected; this is a reliability and availability bug, not an attacker-triggered one. No CVSS score or CWE is in the record.

Who can reach it

No attacker path in the record. Reached by a hardware or firmware-reported PCIe error on a POWER host - a PHB-level error, or a device whose driver is not EEH sensitive. Local or remote reachability is not established.

What to do

Update to a stable kernel that removes the redundant lock/unlock in eeh_rmv_device() and reboot the POWER nodes. The record provides stable commits only, no fixed version numbers. Nodes on x86 or arm64 need no action.

References

Related entries

All Kernel, userspace & hypervisor entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.