Database/Control plane, storage & DevOps
AMD PCIe link handling (memory buffer bounds): A guest VM can drive the PCIe link into an out-of-bounds condition
Impact
A guest VM can drive the PCIe link into an out-of-bounds condition and deny service to the entire host. On an AI node the PCIe fabric is the load-bearing structure - GPUs, NVMe scratch, RDMA NICs all hang off it - so a link-level fault does not degrade one tenant, it takes the box down and kills every job on it. This is a guest-to-host availability break reachable from a normal VM, which is a materially different risk class from the ring 0 firmware issues elsewhere in this set.
Who can reach it
Attacker with access to a guest virtual machine - an ordinary paying tenant. Network-adjacent attack vector per AMD's scoring, low privilege required.
What to do
Firmware update per AMD-SB-4013 from the OEM; BIOS flash and reboot. In the meantime the practical control is blast-radius management rather than prevention: do not co-locate high-value long-running training jobs with untrusted short-lived tenants on the same PCIe complex, and make sure checkpointing intervals assume the node can vanish. Note this bulletin targets client and embedded platform audits, so confirm applicability to your specific EPYC or Instinct SKUs with your vendor before planning a fleet-wide flash.
References
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.