Database/Kernel, userspace & hypervisor
Linux kernel net/rds: a missing barrier in release_in_xmit() loses the wake-up and strands the RDS shutdown worker
Impact
release_in_xmit() cleared RDS_IN_XMIT with clear_bit_unlock() and then read the wait queue with a plain waitqueue_active(). clear_bit_unlock() is release-only and does not order that later load, while the waiter adds itself to the queue and then tests the bit - the classic store-buffering pattern, so the releasing CPU can see an empty queue while the waiter still sees the bit set and never wakes. The waiters are rds_conn_shutdown() and rds_tcp_reset_callbacks(), both in uninterruptible wait_event() with no timeout, so a lost wake-up strands the shutdown worker on its single-threaded workqueue until another sender releases the bit - and on a connection being torn down because it failed, there may never be another sender. That blocks the whole RDS shutdown workqueue: on a cluster node the interconnect teardown hangs with a task in unkillable D state, and clearing it in practice means rebooting. The barrier was present until it was folded into clear_bit_unlock(); the refill counterpart in net/rds/ib_recv.c still carries its smp_mb__after_atomic().
Who can reach it
Local/on-fabric race on a node running the rds module: needs a connection teardown concurrent with a sender releasing RDS_IN_XMIT. No authentication is involved and no tenant-facing trigger is described.
What to do
Update to a stable kernel that uses wq_has_sleeper() - waitqueue_active() preceded by the required full barrier - in release_in_xmit(). Requires a node reboot, and a node already stuck in this state needs one anyway since the waiter is uninterruptible. Nodes that do not need RDS can unload or blacklist the rds modules instead.
References
Related entries
- Linux kernel BPF: a BPF_PSEUDO_FUNC load of the main program is never relocated, leaving a call to a bogus addressCVE-2026-98075 · Linux kernel BPF verifier (BPF_PSEUDO_FUNC reference to the main program)Unscored
- Linux kernel mpt3sas: NUMA_NO_NODE from dev_to_node() causes an out-of-bounds node_to_cpumask_map readCVE-2026-98088 · Linux kernel mpt3sas (_base_assign_reply_queues() NUMA node lookup)Unscored
- Linux kernel mpi3mr: error path in mpi3mr_sas_port_add() leaks a target device referenceCVE-2026-98128 · Linux kernel mpi3mr (target device refcount leak in mpi3mr_sas_port_add())Unscored
- Linux kernel mpi3mr: NULL dereference and sas_port leak when SAS port allocation failsCVE-2026-98129 · Linux kernel mpi3mr (Broadcom tri-mode SAS/SATA/NVMe HBA driver)Unscored
- Linux cgroup: task iterator can resurrect a zero-refcount dying task, giving a use-after-freeCVE-2026-98163 · Linux kernel cgroup task iterator (css_task_iter_next over dying_tasks)Unscored
- Linux KVM x86/mmu: write tracking checked in one address space only, reaching a kernel BUGCVE-2026-98164 · Linux kernel KVM x86/mmu (kvm_gfn_is_write_tracked across address spaces)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.