GPU VulnDB

Database/Firmware, BMC & network fabric

Linux kernel RDMA/counter: failed QP bind leaks the port counter count and wedges counter mode

UnscoredCVE-2026-97477Firmware, BMC & network fabriccurated

Impact

When __rdma_counter_bind_qp() fails, alloc_and_bind() frees the counter without decrementing port_counter->num_counters, and the only code that decrements it is unreachable for a counter that was never bound. Each failure therefore permanently inflates the count, and once inflated the port can no longer be switched to AUTO mode (__counter_set_mode() returns -EBUSY) and the MANUAL-to-NONE auto-revert never happens. For an InfiniBand or RoCE fabric this breaks per-QP statistics on the affected port - the telemetry operators use to find a bad link or a noisy tenant across a GPU training fabric - and the state does not recover without a reboot. No memory corruption, no data disclosure; the damage is to fabric observability and it requires the privilege to manage RDMA counters in the first place.

Who can reach it

Local, privileged: a user or agent able to bind RDMA per-QP counters through the RDMA netlink/configfs interface, typically root on the node or a fabric management daemon. Not reachable from an unprivileged tenant pod.

What to do

Take the stable kernel fix (adds an err_bind label that decrements num_counters and mirrors the MANUAL-to-NONE revert) and reboot the node on the normal kernel-update cadence. The leaked count cannot be reset at runtime, so a node already in this state needs a reboot regardless of the patch.

References

Related entries

All Firmware, BMC & network fabric entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.