GPU VulnDB

Database/Firmware, BMC & network fabric

Linux kernel RDMA core: port speed query can call ethtool on an unregistering netdev

CVSS 7.0CVE-2026-98359Firmware, BMC & network fabriccurated

Impact

ib_get_eth_speed() took a reference on the backing net_device and then invoked its ethtool callback without checking the registration state, so an asynchronous RDMA port query could run against a netdev after NETDEV_UNREGISTER and ndo_uninit had already completed. The record rates it local, low-privilege, high complexity, with confidentiality, integrity and availability impact - in practice a use-after-teardown window hit by racing a port query against RoCE interface removal. On a GPU node this is the RoCE/InfiniBand path that carries NCCL traffic, so the blast radius is a kernel fault on a node that is in the middle of fabric reconfiguration rather than a tenant-crossing read. The race needs an unregister to be in flight, which is an operator action (driver unload, interface teardown, SR-IOV VF removal), not something a tenant can trigger on demand.

Who can reach it

Local user able to issue RDMA port queries (an ioctl/sysfs query on an ib device, no special privilege beyond device access) racing a netdev unregister. Not reachable from the network.

What to do

Pick up the stable kernel commits listed in the record (the fix checks NETREG_REGISTERED under RTNL and returns -ENODEV) and reboot each node once the patched kernel is installed. There is no runtime mitigation; until then, avoid tearing down RoCE interfaces while RDMA clients are active. Rolling this out means draining and rebooting every GPU node, so most operators will fold it into the next scheduled kernel window rather than open one for this alone.

References

Related entries

All Firmware, BMC & network fabric entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.