AMD RCCL: unvalidated peer input lets a compromised rank dereference an attacker-controlled pointer
Impact
RCCL is the collective communication layer for multi-GPU and multi-node training on AMD Instinct parts, the ROCm equivalent of NCCL. A rank that has already been compromised, or an attacker positioned on the network RCCL communicates over, can feed input that causes another rank to dereference a pointer it controls, which AMD says may reach remote code execution. In a shared Instinct cluster this is a lateral-movement primitive: one tenant's compromised training process, or anyone who can speak to the collectives traffic on the fabric, reaches code execution inside another participant's process on another node. AMD lists MI210 through MI355X, so this covers the whole current Instinct datacenter line. Attack complexity is rated high and some privilege is needed, which argues for a scheduled fix rather than an emergency one.
Who can reach it
A compromised peer rank inside a running collectives job, or a network-adjacent attacker able to reach the RCCL transport. Low privilege is required (participation in or reachability of the job's communication path), no user interaction.
What to do
Update RCCL to the fixed version named in AMD bulletin AMD-SB-6033 (ship it with your ROCm or container image refresh). Running jobs link the library, so the fix only takes effect for jobs started afterwards - plan to drain and relaunch long-running training jobs rather than expecting a live patch. No GPU firmware flash or node reboot is indicated. Separately, keep collectives traffic on an isolated fabric that tenants cannot address directly.
References
Related entries
- NVIDIA Windows GPU driver: out-of-bounds read with no privileges required, leaking data and crashing the driverCVE-2026-47576 · NVIDIA GPU Display Driver for Windows (kernel module)High
- Linux amdgpu: unbounded FRU PIA TLV walk reads out of bounds on malformed EEPROM dataCVE-2026-97428 · Linux kernel drm/amdgpu (FRU EEPROM PIA/TLV parser)High
- DGX H100 BMC (IPMI): Credential exposureCVE-2023-25531 · DGX H100 BMC (IPMI)High
- NVIDIA License System - Delegated Licensing Service (DLS): An unauthorised action against the DLS reaches partialCVE-2024-0122 · NVIDIA License System - Delegated Licensing Service (DLS)High
- Container Toolkit / GPU Operator: Container escape to host root (insufficient input validation)CVE-2024-0135 · Container Toolkit / GPU OperatorHigh
- Container Toolkit / GPU Operator: Container escape to host (insufficient input validation)CVE-2024-0136 · Container Toolkit / GPU OperatorHigh
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.