GPU VulnDB

Database/Firmware, BMC & network fabric

Linux kernel RDMA/ucma: race between concurrent multicast join and leave frees in-use state

UnscoredCVE-2026-98253Firmware, BMC & network fabriccurated

Impact

Two concurrent JOIN_MCAST ioctls on the same address through /dev/infiniband/rdma_cm can insert a second CMA multicast entry before the first thread's leave-by-address runs, so leave cancels the wrong work and an older RoCE worker dereferences a ucma_multicast the first thread already freed. The result is a use-after-free reachable from an unprivileged process that holds the rdma_cm character device - on GPU nodes that is any tenant permitted to use RoCE or RDMA verbs from a pod. Consequence is a kernel crash, and potentially corruption of adjacent kernel state, on a node that is carrying fabric traffic for other tenants.

Who can reach it

Local user or container with access to /dev/infiniband/rdma_cm (typical where RDMA/RoCE is exposed to tenant pods). No network authentication needed; no remote path.

What to do

Take the stable-tree fix that holds ctx->mutex from rdma_join_multicast() through copy_to_user() and through the -EFAULT leave path. Shipping it means a kernel update and a reboot of each affected node, so drain GPU workloads first. Until then, restrict /dev/infiniband/rdma_cm exposure to tenant containers. No vendor advisory or CVSS score is published in the record.

References

Related entries

All Firmware, BMC & network fabric entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.